Method for obtaining image recognition model, image recognition method, device and medium

By performing feature extraction and clustering of sample images without category labels, combined with the initial model of sample images with category labels, the problem of how to make full use of sample images without category labels is solved, and the effect of high recognition accuracy and efficient training is achieved.

CN114283310BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110984013.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-25
Publication Date
2025-06-27
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

How to make full use of sample images without category labels to obtain image recognition models and improve recognition accuracy.

Method used

The initial model is obtained by training using the first sample image with the category label, the feature vector of the second sample image without the category label is extracted, and clustered to obtain the clustering result. A second model is obtained based on the clustering result, which is used to identify the category of the input image.

Benefits of technology

Effectively utilize sample images without category labels, improve the generalization and representation capabilities of the image recognition model, thereby improving the recognition accuracy, shortening training time, and saving processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283310B_ABST
    Figure CN114283310B_ABST
Patent Text Reader

Abstract

The present application discloses a method for obtaining an image recognition model, an image recognition method, an apparatus, and a medium, belonging to the field of artificial intelligence technology. The method includes: obtaining a first sample image and a second sample image, where the first sample image is an image with a class label, and the second sample image is an image without a class label. Training a first model based on the first sample image, extracting a feature vector of the second sample image through the first model, and clustering the feature vector to obtain a clustering result. Training a second model based on the clustering result, where the second model is used to identify the class to which the input image belongs. By training the second model based on the clustering result, the present application not only makes full use of the sample images without class labels, but also enables the second model to learn the common characteristics of each second sample image belonging to the same class, improving the recognition accuracy of the second model. The clustering process also saves the processing resources required for training and improves the training efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a method for obtaining an image recognition model, an image recognition method, an apparatus, and a medium. Background Art

[0002] With the development of artificial intelligence technology, the number of sample images in the dataset is also increasing. In the dataset, compared with the sample images with class labels, the number of sample images without class labels is larger. Among them, the class label is used to indicate the class to which the content recorded in the sample image belongs. How to make full use of the sample images without class labels to obtain an image recognition model has become an urgent problem to be solved. Summary of the Invention

[0003] Embodiments of the present application provide a method for obtaining an image recognition model, an image recognition method, an apparatus, and a medium, so as to make full use of the sample images without class labels to obtain an image recognition model and make the image recognition model have a high recognition accuracy. The technical solutions are as follows:

[0004] On the one hand, a method for obtaining an image recognition model is provided, and the method includes:

[0005] Obtain a first sample image and a second sample image, where the first sample image is an image with a class label, and the second sample image is an image without a class label;

[0006] Train a first model based on the first sample image, extract the feature vectors of the second sample image through the first model, and cluster the feature vectors to obtain a clustering result;

[0007] Train a second model based on the clustering result, where the second model is used to identify the class to which the input image belongs.

[0008] On the one hand, an image recognition method is provided, and the method includes:

[0009] Obtain an image to be recognized, input the image into at least two image recognition models respectively to obtain sub-results output by the at least two image recognition models, any one of the image recognition models outputs at least two sub-results, any one of the sub-results corresponds to a class, and any one of the sub-results is used to indicate the probability that the image belongs to the corresponding class. Any one of the image recognition models is trained based on a clustering result, and the clustering result is obtained by clustering the feature vectors of the second sample image extracted by an initial model. The initial model is trained based on the first sample image, the first sample image is an image with a class label, and the second sample image is an image without a class label;

[0010] Weighted sum is performed on sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted sum values;

[0011] Identify the category corresponding to the weighted sum value with the largest indicated probability as the category to which the image belongs.

[0012] On the one hand, an acquisition device for an image recognition model is provided. The device includes:

[0013] An acquisition module, configured to acquire a first sample image and a second sample image, where the first sample image is an image with a category label, and the second sample image is an image without a category label;

[0014] A training module, configured to train a first model based on the first sample image;

[0015] A clustering module, configured to extract feature vectors of the second sample image through the first model, cluster the feature vectors to obtain a clustering result;

[0016] The training module is further configured to train a second model based on the clustering result, and the second model is used to identify the category to which an input image belongs.

[0017] In an exemplary embodiment, the training module is configured to acquire a trained third model, fine-tune the third model based on the first sample image to obtain the first model; fine-tune the third model based on the clustering result to obtain a fine-tuned third model, and obtain the second model based on the fine-tuned third model.

[0018] In an exemplary embodiment, the training module is configured to input the second sample image into the first model to obtain a category label generated by the first model for the second sample image. Any second sample image corresponds to a target sub-result, and the target sub-result is used to indicate the probability that the any second sample image belongs to the category corresponding to the category label; sort the second sample images with the same category label based on the target sub-result to obtain a sample image sequence for each category corresponding to the category label; train the fine-tuned third model based on the sample image sequence for each category corresponding to the category label to obtain the second model.

[0019] In an exemplary embodiment, the training module is configured to, in the target sub-result, filter out target sub-results whose indicated probability is not less than a probability threshold; sort the second sample images with the same category label based on the filtered target sub-results to obtain a sample image sequence for each category corresponding to the category label.

[0020] In an exemplary embodiment, the training module is configured to, for any category corresponding to a category label, obtain at least two sample image subsets of the category corresponding to the any category label from a sample image sequence of the category corresponding to the any category label, where the number of second sample images included in different sample image subsets is different; for any category corresponding to a category label, in the order of gradually changing the number of second sample images, successively train the fine-tuned third model based on each sample image subset of the category corresponding to the any category label to obtain the second model.

[0021] In an exemplary embodiment, the training module is configured to obtain an image set corresponding to each sample image in the first sample image and the second sample image, where the image set corresponding to any sample image includes a global image and a local image obtained based on the any sample image; for any sample image, input the global image included in the image set corresponding to the any sample image into a fourth model to obtain a first output result, input the global image and the local image included in the image set corresponding to the any sample image into a fifth model to obtain a second output result, and determine the cross-entropy loss between the first output result and the second output result; update the fifth model based on the cross-entropy loss to obtain an updated fifth model, and obtain the third model based on the updated fifth model.

[0022] In an exemplary embodiment, the training module is configured to, in response to the processing resources meeting the conditions, update the fourth model based on the updated fifth model to obtain the third model.

[0023] In an exemplary embodiment, the training module is configured to, in response to the processing resources not meeting the conditions, use the updated fifth model as the third model.

[0024] On the one hand, an image recognition device is provided, and the device includes:

[0025] An acquisition module, configured to acquire an image to be recognized, input the image into at least two image recognition models respectively to obtain sub-results output by the at least two image recognition models, where any image recognition model outputs at least two sub-results, any sub-result corresponds to a category, the any sub-result is used to indicate the probability that the image belongs to the corresponding category, the any image recognition model is trained based on a clustering result, the clustering result is obtained by clustering the feature vectors of the second sample images extracted by an initial model, the initial model is trained based on first sample images, the first sample images are images with category labels, and the second sample images are images without category labels;

[0026] A weighted summation module, configured to perform weighted summation on sub-results output by different image recognition models and corresponding to the same category, to obtain at least two weighted summation values;

[0027] An identification module, configured to identify the category corresponding to the weighted summation value with the highest indicated probability as the category to which the image belongs.

[0028] In an exemplary embodiment, the weighted summation module is further configured to determine the accuracy values of each image recognition model, where the accuracy value of any image recognition model is used to indicate the recognition accuracy of the any image recognition model; calculate the sum of the accuracy values of each image recognition model; for any image recognition model, calculate the ratio of the accuracy value of the any image recognition model to the sum of the accuracy values of each image recognition model, and determine the ratio as the weight of at least two sub-results output by the any image recognition model.

[0029] In an exemplary embodiment, the apparatus further includes: an update module, configured to determine, for at least two sub-results output by any image recognition model, the magnification factor corresponding to each sub-result in the at least two sub-results, where for any sub-result, the greater the probability indicated by the any sub-result, the greater the magnification factor corresponding to the any sub-result; update each sub-result according to the magnification factor to obtain updated sub-results;

[0030] The weighted summation module is configured to perform weighted summation on the updated sub-results output by different image recognition models and corresponding to the same category, to obtain the at least two weighted summation values.

[0031] On the one hand, an electronic device is provided, where the electronic device includes a memory and a processor; at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor, so that the electronic device implements the method for obtaining an image recognition model or the image recognition method provided in any exemplary embodiment of the present application.

[0032] On the one hand, a computer-readable storage medium is provided, where at least one instruction is stored in the computer-readable storage medium, and the instruction is loaded and executed by a processor, so that a computer implements the method for obtaining an image recognition model or the image recognition method provided in any exemplary embodiment of the present application.

[0033] On the other hand, a computer program or a computer program product is provided, where the computer program or the computer program product includes: computer instructions, and when the computer instructions are executed by a computer, the computer implements the method for obtaining an image recognition model or the image recognition method provided in any exemplary embodiment of the present application.

[0034] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0035] In this embodiment, a first model is trained using first sample images with class labels. The first model extracts features from second sample images without class labels, and clustering is performed based on the extracted feature vectors to obtain a clustering result. Then, a second model is trained based on the clustering result. Therefore, not only are the second sample images without class labels fully utilized, enabling the second model to have strong generalization ability, but also the second model can learn the common characteristics of each second sample image belonging to the same class, so that the second model has strong representation ability, thereby improving the recognition accuracy of the second model. Moreover, the clustering process is also beneficial for shortening the training duration, saving the processing resources required for training, and improving the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 It is a schematic diagram of the implementation environment provided by the embodiments of the present application;

[0038] Figure 2 It is a flowchart of the method for obtaining an image recognition model provided by the embodiments of the present application;

[0039] Figure 3 It is a schematic diagram of the clustering self-supervised training process provided by the embodiments of the present application;

[0040] Figure 4 It is a schematic diagram of the self-supervised training process provided by the embodiments of the present application;

[0041] Figure 5 It is a schematic diagram of the pseudo-label training process provided by the embodiments of the present application;

[0042] Figure 6 It is a flowchart of the image recognition method provided by the embodiments of the present application;

[0043] Figure 7 It is a schematic diagram of model fusion provided by the embodiments of the present application;

[0044] Figure 8 It is a schematic flowchart of image recognition provided by the embodiments of the present application;

[0045] Figure 9It is a schematic structural diagram of an acquisition device for an image recognition model provided by an embodiment of the present application;

[0046] Figure 10 It is a schematic structural diagram of an image recognition device provided by an embodiment of the present application;

[0047] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0049] The embodiments of the present application provide an image recognition model acquisition method and an image recognition method, and the above methods can be applied to, for example, Figure 1 the implementation environment shown. Figure 1 In, it includes at least one electronic device 11 and a server 12. The electronic device 11 can be communicatively connected to the server 12 to download images that need to be used from the server 12.

[0050] Among them, the electronic device 11 can be any electronic product that can perform human-computer interaction with a user in one or more ways such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device. For example, a PC (Personal Computer), mobile phone, smart phone, PDA (Personal Digital Assistant), wearable device, pocket PC (Pocket PC), tablet computer, intelligent vehicle machine, smart TV, smart speaker, etc.

[0051] The server 12 can be a single server, a server cluster composed of multiple servers, or a cloud computing service center.

[0052] Those skilled in the art should understand that the above electronic device 11 and server 12 are only examples. Other existing or future possible electronic devices or servers that are applicable to the present application should also be included within the protection scope of the present application and are hereby incorporated by reference.

[0053] Based on the above Figure 1 shown implementation environment, refer to Figure 2 The embodiments of the present application provide an image recognition model acquisition method, and this method can be applied to Figure 1 the electronic device shown. As Figure 2 shown, this method includes the following steps.

[0054] 201. Obtain a first sample image and a second sample image. The first sample image is an image with a class label, and the second sample image is an image without a class label.

[0055] Among them, the class label of an image is used to indicate the class to which the image belongs. The class to which the image belongs is also the class of the content recorded in the image, and the class label corresponds to the class one by one. Exemplarily, in this embodiment, the first sample image and the second sample image are obtained from a data set. This embodiment does not limit the data set, and the data set includes but is not limited to the public data set and the private data set in FGVC (Fine-Grained Visual Categorization) 8. Alternatively, in this embodiment, images are collected. Then, among the collected images, class labels are generated for a part of the images to be used as the first sample images. The other part of the images do not generate class labels and are used as the second sample images. Exemplarily, the ways to generate class labels include: generating class labels by manual annotation, or outputting class labels through a trained image classification model. This embodiment does not limit the way to generate class labels.

[0056] In this embodiment, the first sample image and the second sample image are used in the model training process. Exemplarily, the numbers of the first sample image and the second sample image are both multiple to ensure the accuracy of the trained model.

[0057] 202. Train a first model based on the first sample image, extract the feature vectors of the second sample image through the first model, and cluster the feature vectors to obtain a clustering result.

[0058] Among them, the first model trained based on the first sample image has the ability of feature extraction. Therefore, the feature vectors of the second sample image can be extracted through the first model. Then, the feature vectors are clustered to obtain a clustering result. Exemplarily, the clustering methods include but are not limited to K-means (K-means Clustering Algorithm). This embodiment does not limit the clustering method. The clustering result includes at least one vector group, and a vector group includes at least one feature vector of the second sample image. Based on at least one vector group, at least one sample image group can be obtained. A sample image group includes the second sample images corresponding to the feature vectors in a vector group. The second sample images included in a sample image group belong to the same class.

[0059] Refer to formula (1). Formula (1) represents the feature extraction process:

[0060] z = M(x), x ∈ Xu (1)

[0061] In formula (1), X u represents the set of second sample images. Since x ∈ X u , thus x in formula (1) represents the second sample image, and z represents the feature vector of the second sample image extracted by the first model M(·).

[0062] Formula (2) represents the process of clustering the feature vectors:

[0063] y x = Kmeans(z)(2)

[0064] In formula (2), y x represents the clustering result.

[0065] Exemplarily, training the first model based on the first sample image includes: obtaining a first initial model, and training the first initial model based on the first sample image to obtain the first model. During the training process, inputting the first sample image into the first initial model to obtain the output result of the first initial model, and the output result is calculated based on the initial model parameters included in the first initial model. Calculating a loss function based on the output result, minimizing the loss function and performing backpropagation of the gradient, thereby updating the initial model parameters included in the first initial model. Then, loop the process of inputting the first sample image into the first initial model and the subsequent calculation process until the training process stops after meeting the termination condition, thereby obtaining the first model. Exemplarily, meeting the termination condition includes: the loss function calculated based on the output result is less than a first threshold, or the difference between the loss functions calculated twice adjacent to each other is less than a second threshold. In this embodiment, the first threshold and the second threshold are not limited, and the first threshold and the second threshold can be set based on experience.

[0066] This embodiment does not limit the type of the first initial model. The first initial model includes but is not limited to: Resnet (Residual Network) model, ViT (Vision Transformer), and Swin (Shifted Windows)-Transformer, etc. The Resnet model is, for example, Resnet101, Resnet 154. The ViT is, for example, ViT-base, ViT-small. The Swin-Transformer is, for example, Swin-Transformer base, Swin-Transformer large.

[0067] Or, in an exemplary embodiment, refer toFigure 3 , a first model is trained based on a first sample image, including: obtaining a trained third model, and fine-tuning the third model based on the first sample image to obtain the first model. Among them, the process of fine-tuning the third model based on the first sample image to obtain the first model is the same as the process of training the first model based on the first initial model described above, and will not be elaborated here. It should be noted that, compared with the process of training the first model based on the first initial model, the number of loops required for the process of fine-tuning the third model based on the first sample image to obtain the first model is less, which not only saves processing resources, but also shortens the time required to obtain the first model and improves the efficiency of obtaining the first model. In this embodiment, the trained third model is, for example, the Resnet model, ViT, and Swin-Transformer in the above examples, and the type of the third model is not limited in this embodiment.

[0068] Exemplarily, the third model is trained in a self-supervised manner in this embodiment. The self-supervised manner includes, but is not limited to: dino (knowledge distillation with no labels), simCLR (a simple frame work for contrastive learning of visiual representations), MoCo (momentum contrast for unsupervised visiualrepresentation learning), etc. In an exemplary embodiment, obtaining the trained third model includes the following steps 2021-2023.

[0069] 2021, obtain the image sets corresponding to the respective sample images in the first sample image and the second sample image. The image set corresponding to any sample image includes a global image and a local image obtained based on any sample image.

[0070] Among them, the set of the first sample images is denoted as X l , the set of the second sample images is denoted as X u , then any sample image x in the first sample image and the second sample image is expressed as x ∈ X u ∪X lExemplarily, after obtaining the first sample image and the second sample image, this embodiment obtains the image set corresponding to each sample image, and stores the sample image and the image set correspondingly. When it is necessary to obtain the trained third model, the image set corresponding to the stored sample image is obtained. Alternatively, in this embodiment, the image set corresponding to the sample image is obtained when it is necessary to obtain the trained third model. This embodiment does not limit the timing of obtaining the image set corresponding to the sample image.

[0071] Next, the methods for obtaining the global image and the local image based on the sample image will be described separately.

[0072] Method for obtaining the global image based on the sample image: Exemplarily, in response to a sample image not meeting the requirements, the global image is intercepted from the sample image. Alternatively, in response to a sample image meeting the requirements, no interception is performed, and the global image is directly obtained based on the sample image. Among them, the requirements to be met are set according to experience or actual needs, and this embodiment does not limit the requirements to be met. For example, the requirements to be met include at least one of shape and resolution, and the shape is, for example, a square. Taking the requirement to be met including shape as an example, in response to the sample image not being a square, a square global image is intercepted from the sample image. In response to a sample image being a square, the global image is directly obtained based on the sample image.

[0073] For a sample image, the number of global images obtained based on the sample image is at least one, and this embodiment does not limit the number of global images. When it is necessary to obtain more than two global images based on a sample image, in response to the sample image not meeting the requirements, interception can be performed at different positions in the sample image to obtain more than two global images. Alternatively, in response to the sample image meeting the requirements, the image is processed in different ways to obtain more than two global images. This embodiment does not limit the processing method, and the processing method is, for example, rotation, inclination, etc.

[0074] Method for obtaining the local image based on the sample image: The local image is intercepted from the sample image. The number of local images is at least one, and the resolution of the local image is less than the resolution of the global image. It can be understood that the content recorded in the global image may include all the content recorded in the local image, may also include a part of all the content recorded in the local image, or may not include the content recorded in the local image. This embodiment does not limit this.

[0075] 2022, see Figure 4, for any sample image, input the global images included in the image set corresponding to the sample image into the fourth model to obtain a first output result. Input the global images and local images included in the image set corresponding to the sample image into the fifth model to obtain a second output result, and determine the cross-entropy loss between the first output result and the second output result.

[0076] Exemplarily, the fourth model and the fifth model are, for example, the Resnet model, ViT, and Swin-Transformer in the above examples. The types of the fourth model and the fifth model are not limited in this embodiment. In some embodiments, the types of the fourth model and the fifth model are the same. For example, both the fourth model and the fifth model are Resnet 101.

[0077] In this embodiment, any one of the first output result and the second output result includes sub-results corresponding to the input image. One image corresponds to at least two sub-results, and the number of sub-results corresponding to one image is the same as the number of categories that the model can recognize. One sub-result corresponds to one category, and one sub-result is used to indicate the probability that the image belongs to the corresponding category. For example, the fourth model can recognize N categories. After inputting a global image into the fourth model, the fourth model outputs N sub-results corresponding to the global image. The first sub-result is used to indicate the probability that the global image belongs to category 1, the second sub-result is used to indicate the probability that the global image belongs to category 2, and so on. The Nth sub-result is used to indicate the probability that the global image belongs to category N. Wherein, N is a positive integer not less than 2.

[0078] Exemplarily, the sub-result is a confidence value (logit), and the value range of the confidence value is from negative infinity to positive infinity. Or, the sub-result is a probability value obtained by normalizing the confidence value, and the value range of the probability value is from 0 to 1. In some embodiments, the confidence value is normalized by the softmax function. The normalization method is not limited in this embodiment. Whether the sub-result is a confidence value or a probability value, the sub-result can indicate the probability that the image belongs to the corresponding category. In some embodiments, the larger the value of the sub-result, the greater the probability indicated by the sub-result.

[0079] After obtaining the first output result and the second output result, determine the cross-entropy loss between the first output result and the second output result. Since the global image is input into the fourth model in this embodiment, the first output result output by the fourth model includes sub-results corresponding to each global image. Since the local image and the global image are both input into the fifth model in this embodiment, the second output result output by the fifth model includes both sub-results corresponding to each global image and sub-results corresponding to each local image. Exemplarily, in this embodiment, the cross-entropy loss is determined based on the sub-results corresponding to each global image in the first output result and the sub-results corresponding to each local image in the second output result.

[0080] It should be noted that for a sample image, the image set included in the sample image includes at least one global image and at least one local image. A cross-entropy loss can be calculated based on the sub-result corresponding to a global image and the sub-result corresponding to a local image. Correspondingly, in the image set corresponding to a sample image, in response to the number of either the global image or the local image being at least two, at least two cross-entropy losses can be calculated.

[0081] In 2023, update the fifth model based on the cross-entropy loss to obtain the updated fifth model, and obtain the third model based on the updated fifth model.

[0082] According to the description in 2022, the number of cross-entropy losses may be one or at least two. Exemplarily, updating the fifth model based on the cross-entropy loss includes: calculating the sum of each cross-entropy loss, minimizing the sum of the cross-entropy losses and performing gradient descent, so as to realize the update of the fifth model and obtain the updated fifth model. Exemplarily, the way of performing gradient descent includes but is not limited to SGD (Stochastic Gradient Descent), and this embodiment does not limit the way of performing gradient descent.

[0083] Among them, the sum of the cross-entropy losses is expressed as the following formula (3):

[0084]

[0085] In formula (3), I is the sum of the cross-entropy losses. x g is the set composed of global images. Since x ∈ x g, so x is the global image. V is the image set corresponding to the sample image. Since x' ∈ V and x' ≠ x, x' is the local image. P1(x) is the sub-result corresponding to a global image in the first output result, P2(x') is the sub-result corresponding to a local image in the second output result, and H(·, ·) is a cross-entropy loss calculated based on the sub-result corresponding to a global image and the sub-result corresponding to a local image.

[0086] After obtaining the updated fifth model, this embodiment further obtains the third model based on the updated fifth model. In an exemplary embodiment, obtaining the third model based on the updated fifth model includes the following two methods.

[0087] Method 1: In response to the processing resources meeting the conditions, update the fourth model based on the updated fifth model to obtain the third model.

[0088] Among them, since the number of model parameters of the fourth model is greater than that of the fifth model, the processing resources required to use the fourth model are also more than those required to use the fifth model. Therefore, in this embodiment, when the processing resources meet the conditions, that is, when the processing resources are sufficient, the fourth model is used. Among them, using the fourth model is also to update the fourth model based on the updated fifth model, so as to obtain the third model.

[0089] Exemplarily, updating the fourth model based on the updated fifth model to obtain the third model includes: updating the fourth model based on the updated fifth model according to the following formula (4) to obtain the third model:

[0090] θ t2 = l * θ t1 + (1 - l) * θ s (4)

[0091] In formula (4), l is a hyperparameter set according to experience, θ s is the model parameter in the updated fifth model, θ t1 is the model parameter in the fourth model, θ t2 is the model parameter in the third model.

[0092] Method 2: In response to the processing resources not meeting the conditions, use the updated fifth model as the third model.

[0093] In response to the processing resources not meeting the conditions, it means that the processing resources are not sufficient, so it is not suitable to use the fourth model. Therefore, in Method 2, the fourth model is no longer used, but the updated fifth model is directly used as the third model.

[0094] 203. Based on the clustering results, a second model is trained. The second model is used to identify the category to which the input image belongs.

[0095] As can be seen from the description in 202, the clustering results include at least one vector group. Based on the at least one vector group, at least one sample image group can be obtained. The content recorded in each second sample image included in a sample image group belongs to the same category. Exemplarily, training the second model based on the clustering results includes: obtaining at least one sample image group based on the clustering results, and training the second initial model based on each sample image group respectively to obtain the second model. The process of training the second initial model can refer to the process of training the first initial model in 202, which will not be elaborated here. Through this training method, the second model can learn the common characteristics of each second sample image belonging to the same category, improving the accuracy of identifying the category to which the image belongs. Among them, the second initial model includes but is not limited to the Resnet model, ViT, and Swin-Transformer in the above examples.

[0096] As can be seen from the description in 202, in some embodiments, the first model is obtained by fine-tuning the third model. Correspondingly, in an exemplary embodiment, training the second model based on the clustering results includes: fine-tuning the third model based on the clustering results to obtain the fine-tuned third model, and obtaining the second model based on the fine-tuned third model. Among them, the method of obtaining the third model can refer to the description in 2021-2023 above. The process of fine-tuning the third model based on the clustering results is the same as the process of training the second model based on the second initial model in the above description, which will not be elaborated here.

[0097] Exemplarily, obtaining the second model based on the fine-tuned third model includes: using the fine-tuned third model as the second model. Or, refer to Figure 3, in this embodiment, the first sample image is used to fine-tune the fine-tuned third model again to obtain the third model after secondary fine-tuning. In response to the third model after secondary fine-tuning meeting the conditions, the third model after secondary fine-tuning is used as the second model. The conditions to be met, for example, are that the recognition accuracy meets the reference threshold, and this embodiment does not limit the reference threshold. Alternatively, in response to the third model after secondary fine-tuning not meeting the conditions, a loop is performed. The loop process includes: using the third model after secondary fine-tuning to extract features and perform clustering on the second sample image to obtain a new clustering result, and then using the new clustering result and the first sample image to fine-tune the third model after secondary fine-tuning to obtain the third model after tertiary fine-tuning. It is determined whether to perform the loop again based on whether the third model after tertiary fine-tuning meets the conditions. And so on, until a model that meets the conditions is obtained and the loop process ends, and the model that meets the conditions is used as the second model. Alternatively, in this embodiment, the fine-tuned third model is further trained by the pseudo-label method to obtain the second model. In an exemplary embodiment, obtaining the second model based on the fine-tuned third model includes the following steps 2031-2033.

[0098] 2031, refer to Figure 5 , input the second sample image into the first model to obtain the class label generated by the first model for the second sample image.

[0099] Among them, after the second sample image is input into the first model, the first model can output at least two sub-results corresponding to the second sample image. According to the description in 2022, one sub-result corresponds to one class, and one sub-result is used to indicate the probability that the second sample image belongs to the class corresponding to the sub-result. In this embodiment, among the at least two sub-results, the sub-result with the highest indicated probability is used as the target sub-result, so one second sample image corresponds to one target sub-result. Correspondingly, the class corresponding to the target sub-result is the class to which the second sample image belongs, and the class label is the class corresponding to the target sub-result. It can be seen that the target sub-result of one second sample image is used to indicate the probability that the second sample image belongs to the class corresponding to the class label.

[0100] For example, after a second sample image is input into the first model, the first model outputs sub-result 1, sub-result 2, and sub-result 3. Sub-result 1 indicates that the probability that the second sample image belongs to class 1 is 0.9, sub-result 2 indicates that the probability that the second sample image belongs to class 2 is 0.5, and sub-result 3 indicates that the probability that the second sample image belongs to class 3 is 0.1. Therefore, sub-result 1 is used as the target sub-result corresponding to the second sample image, and the class 1 corresponding to the target sub-result (sub-result 1) is used as the class to which the second sample image belongs, and the class label of the second sample image is class 1.

[0101] In 2032, the second sample images with the same class label are sorted based on the target sub-results to obtain a sequence of sample images for each class corresponding to the class label.

[0102] One class includes at least one second sample image, and each of the second sample images included in one class has the same class label. Since one second sample image corresponds to one target sub-result, for one class, the second sample images included in this class can be sorted based on the target sub-results to obtain the sequence of sample images corresponding to this class. Exemplarily, in one class, the second sample images included in this class are sorted in descending order of the probability indicated by the target sub-results to obtain the sequence of sample images corresponding to this class. For example, the class labels of the second sample image 1, the second sample image 2, and the second sample image 3 are all class 1, and the probability indicated by the target sub-result corresponding to the second sample image 1 is 0.8, the probability indicated by the target sub-result corresponding to the second sample image 2 is 0.7, and the probability indicated by the target sub-result corresponding to the second sample image 3 is 0.6. Then, the three second sample images are sorted in the order of the second sample image 1, the second sample image 2, and the second sample image 3 to obtain the sequence of sample images corresponding to class 1. It can be understood that the number of second sample images included in a sequence of sample images in this embodiment is not limited.

[0103] In an exemplary embodiment, sorting the second sample images with the same class label based on the target sub-results to obtain a sequence of sample images for each class corresponding to the class label includes: screening out the target sub-results in the target sub-results whose indicated probability is not less than the probability threshold. Sorting the second sample images with the same class label based on the screened target sub-results to obtain a sequence of sample images for each class corresponding to the class label.

[0104] Among them, in response to the probability indicated by a target sub-result being less than the probability threshold, it means that the probability of the second sample image corresponding to this target sub-result belonging to this class is relatively small. Exemplarily, the probability threshold is not limited in this embodiment, and the probability threshold is, for example, 0.2. Therefore, for this class, the second sample image corresponding to this target sub-result belongs to the noise image. If such a noise image is used to train the fine-tuned third model subsequently, it may affect the recognition accuracy of the trained second model. Therefore, the target sub-results whose indicated probability is less than the probability threshold are deleted in this embodiment. Then, the second sample images are sorted based on the screened target sub-results. The method of sorting the sample images based on the screened target sub-results can refer to the description in 2032 above, and will not be elaborated here. In addition, the probability threshold is not limited in this embodiment.

[0105] Of course, the process of deleting the target sub-results whose indicated probability is less than the probability threshold described above is only an example and is not used to limit this embodiment. Exemplarily, this embodiment may also not delete the target sub-results whose indicated probability is less than the probability threshold. Accordingly, the second sample images in the subsequent second sample images used to train the fine-tuned third model also include noise images, and the fine-tuned third model performs a noisy learning process to obtain the second model.

[0106] In 2033, the fine-tuned third model is trained based on the sample image sequences of the categories corresponding to the respective category labels to obtain the second model.

[0107] Exemplarily, training the fine-tuned third model based on the sample image sequences corresponding to the respective categories includes: for one category, the second sample images included in the sample image sequence corresponding to the category are sequentially input into the fine-tuned third model, so as to implement the training of the fine-tuned third model to obtain the second model.

[0108] Or, in an exemplary embodiment, refer to Figure 5 , for the category corresponding to any category label, at least two sample image subsets of the category corresponding to any category label are obtained from the sample image sequence of the category corresponding to any category label, and the number of second sample images included in different sample image subsets is different. For the category corresponding to any category label, in the order of gradually changing the number of second sample images, the fine-tuned third model is sequentially trained based on the respective sample image subsets of the category corresponding to any category label to obtain the second model.

[0109] Exemplarily, in response to the probabilities indicated by the target sub-results corresponding to the respective second sample images in the sample image sequence corresponding to a category decreasing in sequence, then one sample image subset includes the first q (top-q) second sample images in the sample image sequence, and the value of q is different in different sample image subsets. Taking a category corresponding to 4 sample image subsets as an example, the values of q in the 4 sample image subsets are (40, 60, 80, 100) respectively, representing that the 4 sample image subsets include 40, 60, 80, and 100 second sample images respectively. It can be understood that the number of sample image subsets and the number of second sample images included in one sample image subset in this embodiment are not limited.

[0110] In some embodiments, the second model is obtained by successively training the fine-tuned third model based on each subset of sample images corresponding to any category in the order of gradual change in the number of second sample images, including: training the fine-tuned third model through each subset of sample images in the order from the smallest to the largest number of second sample images included, to obtain the second model. That is to say, among the subsets of sample images, the fine-tuned third model is first trained with the subset of sample images including the fewest second sample images, and finally trained with the subset of sample images including the most second sample images. Taking the 4 subsets of sample images in the above example as an example, the fine-tuned third model is successively trained with the subsets of sample images including 40, 60, 80, and 100 second sample images, so as to obtain the second model.

[0111] It should be noted that when the number of second sample images included in a category is greater than the number threshold, this category corresponds to at least two subsets of sample images. And when the number of second sample images included in a category is less than the number threshold, exemplarily, this category includes only one subset of sample images, and all the second sample images in this category are included in this subset of sample images.

[0112] In summary, in this embodiment, the first model is trained using the first sample images with category labels, the second sample images without category labels are feature-extracted by the first model, and clustering is performed based on the extracted feature vectors to obtain a clustering result. Then, the second model is trained based on the clustering result. Therefore, not only the second sample images without category labels are fully utilized, making the second model have strong generalization ability, but also the second model can learn the common characteristics of each second sample image belonging to the same category, so that the second model has strong representation ability, thereby improving the recognition accuracy of the second model. And the clustering process is also beneficial to shortening the training duration, saving the processing resources required for training, and improving the training efficiency.

[0113] The second model trained in this embodiment can be used to identify the category to which an image belongs, and thus can be used to complete tasks involving the image recognition process. Tasks involving the image recognition process are, for example, extraction of high-quality videos, filtering of low-quality videos, etc. In addition, the second model trained in this embodiment can also be combined with other algorithms and applied to the bottom layer of various algorithms. For example, taking the second model trained in this embodiment as an initial model, and using other algorithms to fine-tune this initial model, so as to train a new model. It can be seen that the second model trained in this embodiment is applicable to multiple scenarios and has strong practicability.

[0114] Based on the above Figure 1 shown implementation environment, see Figure 6, an embodiment of the present application further provides an image recognition method, which can be applied to the Figure 1 electronic device shown in Figure 6 . As shown in

[0115] , the method includes the following steps.

[0116] Among them, an image recognition model outputs at least two sub-results corresponding to the image, one sub-result corresponds to one category, and any sub-result is used to indicate the probability that the image belongs to the corresponding category. For the description of the sub-results, reference can be made to 2022 above, and details will not be elaborated here.

[0117] It should be noted that an image recognition model is trained based on the clustering result, and the clustering result is obtained by clustering the feature vectors of the second sample images extracted by the initial model. The initial model is trained based on the first sample images, the first sample images are images with category labels, and the second sample images are images without category labels. Exemplarily, the initial model is the first model in 201-203 above, and the image recognition model is the second model in 201-203 above. Exemplarily, at least two image recognition models are trained in the manner described in 201-203 above, and the types of different image recognition models are different. For example, a total of 6 image recognition models are trained in this embodiment, and the types of the 6 image recognition models are: Resnet101, Resnet154, ViT-base, Vit-small, Swin-transformer-base, and Swin-transformer-large.

[0118] In an exemplary embodiment, referring to Figure 7 , after obtaining the sub-results output by at least two image recognition models, the method further includes: for at least two sub-results output by any image recognition model, determining the magnification factor corresponding to each sub-result among the at least two sub-results. For any sub-result, the greater the probability indicated by any sub-result, the greater the magnification factor corresponding to any sub-result. Update the sub-results output by at least two image recognition models according to the magnification factor to obtain the updated sub-results. By updating the sub-results, the gap between different sub-results can be increased, which is beneficial to improving the subsequent recognition accuracy.

[0119] Exemplarily, for at least two sub-results output by an image recognition model, sort each sub-result in descending order according to the indicated probability, select the first reference number of sub-results in the sequence, and determine the magnification factor corresponding to each sub-result among the first reference number of sub-results. Exemplarily, denote the reference number as K. For the k-th sub-result (k ∈ K) among the first K sub-results, the magnification factor corresponding to this sub-result is (K + 1 - k). Correspondingly, in this embodiment, the sub-results are updated according to the following formula (5) to obtain the updated sub-results:

[0120] logit′ i,k =(K + 1 - k)*logit i,k (5)

[0121] In formula (5), logit i,k is the k-th sub-result in the i-th model, (K + 1 - k) is the magnification factor corresponding to the h-th sub-result in the i-th model, and logit′ i,k is the k-th updated sub-result in the i-th model.

[0122] For example, when the value of K is 5, an image recognition model outputs 10 sub-results for the input image. In this embodiment, the top-5 sub-results among the 10 sub-results are selected, and the probabilities indicated by the 5 sub-results are 0.9, 0.8, 0.7, 0.6, and 0.5 in sequence. Among them, the magnification factor corresponding to the first sub-result is 5, so the updated first sub-result is 5 * 0.9 = 4.5. The magnification factor corresponding to the second sub-result is 4, so the updated second sub-result is 4 * 0.8 = 3.2. The magnification factor corresponding to the third sub-result is 3, so the updated third sub-result is 3 * 0.7 = 2.1. The magnification factor corresponding to the fourth sub-result is 2, so the updated fourth sub-result is 2 * 0.6 = 1.2. The magnification factor corresponding to the fifth sub-result is 1, so the updated fifth sub-result is 1 * 0.5 = 0.5. Therefore, the 5 updated sub-results are 4.5, 3.2, 2.1, 1.2, and 0.5 in sequence. Compared with the sub-results 0.9, 0.8, 0.7, 0.6, and 0.5, the gap between the updated sub-results is larger, which is beneficial to increasing the recognition accuracy of the subsequent image category.

[0123] 602, see Figure 7 , perform weighted summation on the sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted summation values.

[0124] It can be understood that before weighted summation of sub-results, it is necessary to determine the weights of each sub-result. Therefore, in an exemplary embodiment, before weighted summation of sub-results output by different image recognition models and corresponding to the same category, the method further includes: determining the accuracy value of each image recognition model, where the accuracy value of any image recognition model is used to indicate the recognition accuracy of any image recognition model; calculating the sum of the accuracy values of each image recognition model; for any image recognition model, calculating the ratio of the accuracy value of any image recognition model to the sum of the accuracy values of each image recognition model, and determining the ratio as the weight of at least two sub-results output by any image recognition model.

[0125] Among them, the weight of an image recognition model is calculated based on the following formula (6):

[0126]

[0127] In formula (6), a i is the weight of the i-th image recognition model, acc i is the accuracy value of the i-th image recognition model, and ACC is the sum of the accuracy values of each image recognition model.

[0128] In this embodiment, the weights of each sub-result output by an image recognition model are the same as the weight of the image recognition model. After determining the weights of each sub-result, weighted summation is performed on sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted summation values. Exemplarily, in each sub-result output by each image recognition model in this embodiment, the first reference number of sub-results are selected. Among the selected sub-results, weighted summation is performed on sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted summation values. The definition of the first reference number of sub-results is as described above and will not be elaborated here.

[0129] Among them, the process of calculating the weighted summation value is expressed as the following formula (7):

[0130]

[0131] Among them, logit mean is the weighted summation value, sum(model) is the total number of image recognition models, and logit′ i,k are sub-results output by different image recognition models and corresponding to the same category.

[0132] Refer to Table 1 below. Taking the number of image recognition models as 6 and the reference number as 5 (that is, selecting the top-5 sub-results output by each image recognition model for the image, and the 5 sub-results correspond to categories 1-5 respectively) as an example, the process of calculating the weighted sum value is described as follows:

[0133] Table 1

[0134]

[0135] Among them, the weighted sum value corresponding to category 1 is calculated according to the following formula. For the weighted sum values corresponding to other categories, they will not be elaborated here one by one.

[0136] (a1logit 1,1 +a2logit 2,1 +a3logit 3,1 +a4logit 4,1 +a5logit 5,1 +a6logit 6,1 ) / 6

[0137] Of course, for the case of updating the sub-results in 601, exemplarily, the sub-results output by different image recognition models and corresponding to the same category are weighted and summed, including: weighting and summing the updated sub-results output by different image recognition models and corresponding to the same category. That is: replacing the logit i,k in the above formula (7) with logit′ i,k calculated based on formula (6). Among them, the method of determining the weights of the updated sub-results refers to the above description and will not be elaborated here.

[0138] 603, refer to Figure 7 , and identify the category corresponding to the weighted sum value with the largest indicated probability as the category to which the image belongs.

[0139] After obtaining at least two weighted sum values, identify the category corresponding to the weighted sum value with the largest indicated probability as the category to which the image belongs. Among them, since the weighted sum value is calculated based on the sub-results corresponding to the same category, the categories corresponding to the respective sub-results used to calculate the weighted sum value are the categories corresponding to the weighted sum value. For example, based on Table 1 above, if the probabilities indicated by the weighted sum values corresponding to categories 1, 2, 3, 4, and 5 decrease in sequence, then category 1 is taken as the category of the image to be recognized in 601.

[0140] In summary, in this embodiment, sub-results of an image to be recognized are respectively output by different image recognition models. Then, the sub-results output by different image recognition models and corresponding to the same category are weighted and summed, and the category to which the image to be recognized belongs is determined based on the probability indicated by the weighted sum value. This embodiment is equivalent to fusing at least two image recognition models, improving the recognition accuracy.

[0141] See Figure 8 , Figure 8 which shows a schematic flowchart of an exemplary image recognition provided by an embodiment of the present application. Among them, after obtaining a first sample image with a category label and a second sample image without a category label in this embodiment, a third model is first obtained through the dino self-supervised process (2021-2023). Then, a clustering self-supervised process (201, 202, and 203) is executed. During the clustering self-supervised process, the third model is fine-tuned based on the first sample image to obtain a first model, the feature vectors of the second sample image are extracted through the first model, the feature vectors are clustered to obtain a clustering result, and the third model is fine-tuned based on the clustering result to obtain a fine-tuned third model. Next, a pseudo-label training process (2031-2033) is executed, where the category label of the second sample image is generated through the first model, so as to generate a subset of sample images corresponding to each category label, and the fine-tuned third model is trained based on the subset of sample images to obtain a second model. In some embodiments, at least two different types of second models are trained through the above dino self-supervised process, clustering self-supervised process, and pseudo-label process, and then a model fusion process (601-603) is executed to facilitate the recognition of the category to which the image belongs.

[0142] Exemplarily, in this embodiment, a Vit-small model with a recognition accuracy of 55.4% is obtained, and a third model is obtained through the dino self-supervised process, and the recognition accuracy of the third model is 66.3%. Then, a fine-tuned third model is obtained through the clustering self-supervised process, and the recognition accuracy of the fine-tuned third model is 67.8%. Next, a second model is obtained through the pseudo-label process, and the recognition accuracy of the second model is increased to 70.1%. Finally, through the model fusion process, the recognition accuracy is further increased by 2%.

[0143] The method embodiments provided in the present application have been described above. The image recognition model trained by the method embodiments provided in this embodiment ranks second in the recognition accuracy of the image category on the public dataset of FGVC8 and ranks third in the recognition accuracy of the image category on the private dataset. Exemplarily, the recognition accuracy of the graphic recognition model is indicated by the top-1 error (error rate). The smaller the top-1 error, the higher the recognition accuracy. The top-1 error rankings of each model are shown in Table 2 below:

[0144] Table 2

[0145]

[0146] The embodiments of the present application provide an acquisition device for an image recognition model. Refer to Figure 9 , the device includes:

[0147] An acquisition module 901, configured to acquire a first sample image and a second sample image, where the first sample image is an image with a category label, and the second sample image is an image without a category label;

[0148] A training module 902, configured to train a first model based on the first sample image;

[0149] A clustering module 903, configured to extract a feature vector of the second sample image through the first model, cluster the feature vector, and obtain a clustering result;

[0150] The training module 902 is further configured to train a second model based on the clustering result, and the second model is used to identify the category to which the input image belongs.

[0151] In an exemplary embodiment, the training module 902 is configured to obtain a trained third model, fine-tune the third model based on the first sample image to obtain a first model; fine-tune the third model based on the clustering result to obtain a fine-tuned third model, and obtain a second model based on the fine-tuned third model.

[0152] In an exemplary embodiment, the training module 902 is configured to input the second sample image into the first model to obtain a category label generated by the first model for the second sample image. Any second sample image corresponds to a target sub-result, and the target sub-result is used to indicate the probability that any second sample image belongs to the category corresponding to the category label; sort the second sample images with the same category label based on the target sub-result to obtain a sample image sequence of the category corresponding to each category label; train the fine-tuned third model based on the sample image sequence of the category corresponding to each category label to obtain a second model.

[0153] In an exemplary embodiment, the training module 902 is configured to screen out target sub-results in the target sub-results whose indicated probability is not less than a probability threshold; and sort the second sample images with the same class label based on the screened-out target sub-results to obtain a sequence of sample images for each class corresponding to each class label.

[0154] In an exemplary embodiment, for any class corresponding to a class label, the training module 902 is configured to obtain at least two subsets of sample images for the class corresponding to any class label from the sequence of sample images for the class corresponding to any class label, and the number of second sample images included in different subsets of sample images is different; for any class corresponding to a class label, train the fine-tuned third model based on each subset of sample images for the class corresponding to any class label in the order of gradually changing the number of second sample images, to obtain a second model.

[0155] In an exemplary embodiment, the training module 902 is configured to obtain the image sets corresponding to each sample image in the first sample image and the second sample image, and the image set corresponding to any sample image includes a global image and a local image obtained based on any sample image; for any sample image, input the global image included in the image set corresponding to any sample image into a fourth model to obtain a first output result, input the global image and the local image included in the image set corresponding to any sample image into a fifth model to obtain a second output result, and determine the cross-entropy loss between the first output result and the second output result; update the fifth model based on the cross-entropy loss to obtain an updated fifth model, and obtain a third model based on the updated fifth model.

[0156] In an exemplary embodiment, the training module 902 is configured to update the fourth model based on the updated fifth model to obtain a third model in response to the processing resources meeting the conditions.

[0157] In an exemplary embodiment, the training module 902 is configured to use the updated fifth model as the third model in response to the processing resources not meeting the conditions.

[0158] In summary, in this embodiment, a first model is trained using the first sample images with class labels, the second sample images without class labels are feature-extracted by the first model, and clustering is performed based on the extracted feature vectors to obtain a clustering result. Then, a second model is trained based on the clustering result. Therefore, not only the second sample images without class labels are fully utilized, making the second model have strong generalization ability, but also the second model can learn the common characteristics of each second sample image belonging to the same class, so that the second model has strong representation ability, thereby improving the recognition accuracy of the second model. And the clustering process is also beneficial to shortening the training duration, saving the processing resources required for training, and improving the training efficiency.

[0159] The embodiments of the present application also provide an image recognition device. Refer to Figure 10 , the device includes:

[0160] An acquisition module 1001, configured to acquire an image to be recognized, input the image into at least two image recognition models respectively, obtain sub-results output by the at least two image recognition models. Any one of the image recognition models outputs at least two sub-results. Any one of the sub-results corresponds to a category, and any one of the sub-results is used to indicate the probability that the image belongs to the corresponding category. Any one of the image recognition models is trained based on a clustering result, and the clustering result is obtained by clustering the feature vectors of the second sample images extracted by the initial model. The initial model is trained based on the first sample images, and the first sample images are images with category labels, and the second sample images are images without category labels;

[0161] A weighted summation module 1002, configured to perform weighted summation on the sub-results output by different image recognition models and corresponding to the same category, to obtain at least two weighted summation values;

[0162] An identification module 1003, configured to identify the category corresponding to the weighted summation value with the largest indicated probability as the category to which the image belongs.

[0163] In an exemplary embodiment, the weighted summation module 1002 is further configured to determine the accuracy values of each image recognition model. The accuracy value of any one of the image recognition models is used to indicate the recognition accuracy of any one of the image recognition models; calculate the sum of the accuracy values of each image recognition model; for any one of the image recognition models, calculate the ratio of the accuracy value of any one of the image recognition models to the sum of the accuracy values of each image recognition model, and determine the ratio as the weight of at least two sub-results output by any one of the image recognition models.

[0164] In an exemplary embodiment, the device further includes: an update module, configured to, for at least two sub-results output by any one of the image recognition models, determine the magnification factor corresponding to each of the at least two sub-results. For any one of the sub-results, the greater the probability indicated by any one of the sub-results, the greater the magnification factor corresponding to any one of the sub-results; update each of the sub-results according to the magnification factor to obtain updated sub-results;

[0165] The weighted summation module 1002 is configured to perform weighted summation on the updated sub-results output by different image recognition models and corresponding to the same category, to obtain at least two weighted summation values.

[0166] In summary, in this embodiment, sub-results of an image to be recognized are respectively output by different image recognition models. Then, the sub-results output by different image recognition models and corresponding to the same category are weighted and summed, and the category to which the image to be recognized belongs is determined based on the probability indicated by the weighted sum value. This embodiment is equivalent to fusing at least two image recognition models, improving the recognition accuracy.

[0167] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.

[0168] See Figure 11 , which shows a schematic structural diagram of an electronic device 1100 provided in an embodiment of the present application. The electronic device 1100 may be a portable mobile electronic device, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The electronic device 1100 may also be referred to by other names such as a user equipment, a portable electronic device, a laptop electronic device, a desktop electronic device, etc.

[0169] Generally, the electronic device 1100 includes: a processor 1101 and a memory 1102.

[0170] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one hardware form selected from the group consisting of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen 1105. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0171] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1101 to implement the method for obtaining an image recognition model or the image recognition method provided in the method embodiments of the present application.

[0172] In some embodiments, the electronic device 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one selected from the group consisting of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1109.

[0173] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0174] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1104 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or Wi-Fi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0175] The display screen 1105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input to the processor 1101 as control signals for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, which is provided on the front panel of the electronic device 1100; in other embodiments, there may be at least two display screens 1105, which are respectively provided on different surfaces of the electronic device 1100 or are in a folding design; in other embodiments, the display screen 1105 may be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 1100. Even, the display screen 1105 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1105 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0176] The camera module 1106 is used to collect images or videos. Optionally, the camera module 1106 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1106 may also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. A two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0177] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0178] The power supply 1109 is used to supply power to each component in the electronic device 1100. The power supply 1109 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 1109 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0179] In some embodiments, the electronic device 1100 further includes one or more sensors 1110. The one or more sensors 1110 include but are not limited to: an acceleration sensor 1111, a gyroscope sensor 1112, a pressure sensor 1113, an optical sensor 1115, and a proximity sensor 1116.

[0180] The acceleration sensor 1111 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the electronic device 1100. For example, the acceleration sensor 1111 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1101 can control the display screen 1105 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1111. The acceleration sensor 1111 can also be used for collecting game or user's motion data.

[0181] The gyroscope sensor 1112 can detect the body direction and rotation angle of the electronic device 1100. The gyroscope sensor 1112 can cooperate with the acceleration sensor 1111 to collect the 3D actions of the user on the electronic device 1100. According to the data collected by the gyroscope sensor 1112, the processor 1101 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0182] The pressure sensor 1113 can be disposed on the side frame of the electronic device 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side frame of the electronic device 1100, it can detect the holding signal of the user on the electronic device 1100, and the processor 1101 performs left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 controls the operable controls on the UI interface according to the pressure operation of the user on the display screen 1105. The operable controls include at least one of the group consisting of button controls, scroll bar controls, icon controls, and menu controls.

[0183] The optical sensor 1115 is used to collect the ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 according to the ambient light intensity collected by the optical sensor 1115. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 11011 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera module 1106 according to the ambient light intensity collected by the optical sensor 1115.

[0184] The proximity sensor 1116, also known as the distance sensor, is usually disposed on the front panel of the electronic device 1100. The proximity sensor 1116 is used to collect the distance between the user and the front of the electronic device 1100. In one embodiment, when the proximity sensor 1116 detects that the distance between the user and the front of the electronic device 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from the lit state to the off state; when the proximity sensor 1116 detects that the distance between the user and the front of the electronic device 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from the off state to the lit state.

[0185] Those skilled in the art can understand that Figure 11 the structure shown in does not constitute a limitation on the electronic device 1100, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0186] The embodiment of the present application provides an electronic device, which includes a memory and a processor; at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor so that the electronic device implements the method for obtaining an image recognition model or the image recognition method provided by any one of the exemplary embodiments of the present application.

[0187] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to enable a computer to implement the method for obtaining an image recognition model or the image recognition method provided in any exemplary embodiment of the present application.

[0188] An embodiment of the present application provides a computer program or a computer program product, which includes: computer instructions. When the computer instructions are executed by a computer, the computer is enabled to implement the method for obtaining an image recognition model or the image recognition method provided in any exemplary embodiment of the present application.

[0189] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here.

[0190] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, or the like.

[0191] The above are only embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for obtaining an image recognition model, characterized in that, The method includes: Obtaining a first sample image and a second sample image, where the first sample image is an image with a class label, and the second sample image is an image without a class label; Obtaining an image set corresponding to each sample image in the first sample image and the second sample image, where the image set corresponding to any sample image includes a global image and a local image obtained based on the any sample image; For any sample image, inputting the global image included in the image set corresponding to the any sample image into a fourth model to obtain a first output result, and inputting the global image and the local image included in the image set corresponding to the any sample image into a fifth model to obtain a second output result. Any one of the first output result and the second output result includes at least two sub-results corresponding to each input image, and the sub-result corresponding to each input image indicates the probability that the corresponding image belongs to a certain class; Determining a cross-entropy loss based on the sub-results corresponding to the global images in the first output result and the sub-results corresponding to the local images in the second output result; Updating the fifth model based on the cross-entropy loss to obtain an updated fifth model, and obtaining a third model based on the updated fifth model; Fine-tuning the third model based on the first sample image to obtain a first model; Extracting a feature vector of the second sample image through the first model, and clustering the feature vector to obtain a clustering result; Fine-tuning the third model based on the clustering result to obtain a fine-tuned third model, and obtaining a second model based on the fine-tuned third model. The second model is used to identify the class to which an input image belongs.

2. The method according to claim 1, characterized in that, The obtaining the second model based on the fine-tuned third model includes: Inputting the second sample image into the first model to obtain a class label generated by the first model for the second sample image. Any second sample image corresponds to a target sub-result, and the target sub-result is used to indicate the probability that the any second sample image belongs to the class corresponding to the class label; Sorting the second sample images with the same class label based on the target sub-result to obtain a sample image sequence for each class corresponding to each class label; Training the fine-tuned third model based on the sample image sequences for each class corresponding to each class label to obtain the second model.

3. The method according to claim 2, wherein The sorting the second sample images with the same class label based on the target sub-result to obtain a sample image sequence for each class corresponding to each class label includes: In the target sub-results, screening out the target sub-results whose indicated probability is not less than a probability threshold; Sorting the second sample images with the same class label based on the screened target sub-results to obtain a sample image sequence for each class corresponding to each class label.

4. The method according to claim 2 or 3, characterized in that, The training the fine-tuned third model based on the sample image sequences for each class corresponding to each class label to obtain the second model includes: For the category corresponding to any category label, obtain at least two sample image subsets of the category corresponding to the any category label from the sample image sequence of the category corresponding to the any category label, and the number of second sample images included in different sample image subsets is different; For the category corresponding to any category label, in the order of gradually changing the number of second sample images, successively train the fine-tuned third model based on each sample image subset of the category corresponding to the any category label to obtain the second model.

5. The method according to any one of claims 1 to 3, characterized in that, The obtaining the third model based on the updated fifth model includes: In response to the processing resources meeting the conditions, update the fourth model based on the updated fifth model to obtain the third model.

6. The method according to any one of claims 1-3, characterized in that, The obtaining the third model based on the updated fifth model includes: In response to the processing resources not meeting the conditions, use the updated fifth model as the third model.

7. An image recognition method, characterized in that, The method includes: Obtain an image to be recognized, input the image into at least two image recognition models respectively to obtain sub-results output by the at least two image recognition models. Any image recognition model outputs at least two sub-results, and any sub-result corresponds to a category. The any sub-result is used to indicate the probability that the image belongs to the corresponding category. The any image recognition model is obtained based on a fine-tuned third model, the fine-tuned third model is obtained by fine-tuning the third model based on a clustering result, the clustering result is obtained by clustering the feature vectors of the second sample images extracted by an initial model, the initial model is obtained by fine-tuning the third model based on first sample images, the third model is obtained based on an updated fifth model, the updated fifth model is obtained by updating the fifth model based on a cross-entropy loss, the cross-entropy loss is determined based on the sub-results corresponding to the global images in the first output result and the sub-results corresponding to the local images in the second output result. The first output result is obtained by inputting the global images included in the image set corresponding to any sample image into a fourth model, the second output result is obtained by inputting the global images and local images included in the image set corresponding to the any sample image into the fifth model. Any output result in the first output result and the second output result includes at least two sub-results corresponding to each input image, and the sub-results corresponding to each input image indicate the probability that the corresponding image belongs to a category. The image set corresponding to any sample image is obtained based on the first sample images and the second sample images. The image set corresponding to any sample image includes the global image and local image obtained based on the any sample image. The first sample image is an image with a category label, and the second sample image is an image without a category label; Perform weighted summation on the sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted summation values; Identify the category corresponding to the weighted summation value with the largest indicated probability as the category to which the image belongs.

8. The method according to claim 7, wherein Before weighted summation of sub-results output by different image recognition models and corresponding to the same category, the method further includes: Determine the accuracy values of each image recognition model, where the accuracy value of any image recognition model is used to indicate the recognition accuracy of the any image recognition model; Calculate the sum of the accuracy values of each image recognition model; For any image recognition model, calculate the ratio of the accuracy value of the any image recognition model to the sum of the accuracy values of each image recognition model, and determine the ratio as the weight of at least two sub-results output by the any image recognition model.

9. The method according to claim 7, wherein After obtaining the sub-results output by at least two image recognition models, the method further includes: For at least two sub-results output by any image recognition model, determine the magnification factor corresponding to each sub-result in the at least two sub-results. For any sub-result, the greater the probability indicated by the any sub-result, the greater the magnification factor corresponding to the any sub-result; Update each sub-result according to the magnification factor to obtain updated sub-results; The weighted summation of sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted summation values includes: Perform weighted summation on the updated sub-results output by different image recognition models and corresponding to the same category to obtain the at least two weighted summation values.

10. An apparatus for obtaining an image recognition model, characterized in that The device includes: An acquisition module, configured to acquire a first sample image and a second sample image, where the first sample image is an image with a category label, and the second sample image is an image without a category label; A training module, configured to acquire an image set corresponding to each sample image in the first sample image and the second sample image. The image set corresponding to any sample image includes a global image and a local image acquired based on the any sample image; for any sample image, input the global image included in the image set corresponding to the any sample image into a fourth model to obtain a first output result, input the global image and the local image included in the image set corresponding to the any sample image into a fifth model to obtain a second output result. Any output result in the first output result and the second output result includes at least two sub-results corresponding to each input image, and the sub-result corresponding to each input image indicates the probability that the corresponding image belongs to a category; based on the sub-results corresponding to each global image in the first output result and the sub-results corresponding to each local image in the second output result, determine the cross-entropy loss; update the fifth model based on the cross-entropy loss to obtain an updated fifth model, and obtain a third model based on the updated fifth model; fine-tune the third model based on the first sample image to obtain a first model; A clustering module, configured to extract a feature vector of the second sample image through the first model, and perform clustering on the feature vector to obtain a clustering result; The training module is further configured to fine-tune the third model based on the clustering result to obtain a fine-tuned third model, and obtain a second model based on the fine-tuned third model, where the second model is used to identify the category to which the input image belongs.

11. The device according to claim 10, characterized in that, The training module is configured to input the second sample image into the first model to obtain a category label generated by the first model for the second sample image, and any second sample image corresponds to a target sub-result, where the target sub-result is used to indicate the probability that the any second sample image belongs to the category corresponding to the category label; Sort the second sample images with the same category label based on the target sub-result to obtain a sample image sequence for each category corresponding to the category label; Train the fine-tuned third model based on the sample image sequences for each category corresponding to the category labels to obtain the second model.

12. The device according to claim 11, characterized in that, The training module is configured to screen out the target sub-results in which the indicated probability is not less than the probability threshold in the target sub-results; sort the second sample images with the same category label based on the screened target sub-results to obtain a sample image sequence for each category corresponding to the category label.

13. The device according to claim 11 or 12, characterized in that, For each category corresponding to any category label, the training module is configured to obtain at least two sample image subsets for the category corresponding to the any category label from the sample image sequence for the category corresponding to the any category label, and the number of second sample images included in different sample image subsets is different; for each category corresponding to any category label, train the fine-tuned third model based on each sample image subset for the category corresponding to the any category label in the order of gradual change of the number of second sample images to obtain the second model.

14. The device according to any one of claims 10 to 12, characterized in that, The training module is configured to update the fourth model based on the updated fifth model to obtain the third model in response to the processing resources meeting the conditions.

15. The device according to any one of claims 10 to 12, characterized in that The training module is configured to use the updated fifth model as the third model in response to the processing resources not meeting the conditions.

16. An image recognition device, characterized in that, The device includes: An acquisition module, configured to acquire an image to be recognized, input the image into at least two image recognition models respectively, and obtain sub-results output by the at least two image recognition models. Any one of the image recognition models outputs at least two sub-results, and any one of the sub-results corresponds to a category. The any one of the sub-results is used to indicate the probability that the image belongs to the corresponding category. Any one of the image recognition models is obtained based on a fine-tuned third model, and the fine-tuned third model is obtained by fine-tuning the third model based on a clustering result. The clustering result is obtained by clustering the feature vectors of the second sample images extracted by an initial model. The initial model is obtained by fine-tuning the third model based on the first sample images. The third model is obtained based on an updated fifth model, and the updated fifth model is obtained by updating the fifth model based on a cross-entropy loss. The cross-entropy loss is determined based on the sub-results corresponding to the global images in the first output result and the sub-results corresponding to the local images in the second output result. The first output result is obtained by inputting the global images included in the image set corresponding to any one of the sample images into a fourth model. The second output result is obtained by inputting the global images and local images included in the image set corresponding to the any one of the sample images into the fifth model. Any one of the first output result and the second output result includes at least two sub-results corresponding to each of the input images. The sub-results corresponding to each of the input images indicate the probability that the corresponding image belongs to a category. The image set corresponding to any one of the sample images is obtained based on the first sample images and the second sample images. The image set corresponding to any one of the sample images includes the global images and local images obtained based on the any one of the sample images. The first sample images are images with category labels, and the second sample images are images without category labels; A weighted summation module, configured to perform weighted summation on the sub-results output by different image recognition models and corresponding to the same category, to obtain at least two weighted summation values; An identification module, configured to identify the category corresponding to the weighted summation value with the largest indicated probability as the category to which the image belongs.

17. The device according to claim 16, characterized in that The weighted summation module is further configured to determine the accuracy values of the respective image recognition models. The accuracy value of any one of the image recognition models is used to indicate the recognition accuracy of the any one of the image recognition models; Calculate the sum of the accuracy values of the respective image recognition models; for any one of the image recognition models, calculate the ratio of the accuracy value of the any one of the image recognition models to the sum of the accuracy values of the respective image recognition models, and determine the ratio as the weight of the at least two sub-results output by the any one of the image recognition models.

18. The apparatus according to claim 16, wherein, The device further includes: an updating module, configured to determine, for at least two sub-results output by any image recognition model, an enlargement multiple corresponding to each of the at least two sub-results, wherein for any sub-result, the greater the probability indicated by the any sub-result, the greater the enlargement multiple corresponding to the any sub-result; update the respective sub-results according to the enlargement multiple to obtain updated sub-results; The weighted summation module is configured to perform weighted summation on the updated sub-results output by different image recognition models and corresponding to the same category to obtain at least two weighted summation values.

19. An electronic device, characterized in that, The electronic device includes a memory and a processor; at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor so that the electronic device implements the method for obtaining an image recognition model according to any one of claims 1-6 or the image recognition method according to any one of claims 7-9.

20. A computer-readable storage medium, characterized in that, At least one instruction is stored in the computer-readable storage medium, and the instruction is loaded and executed by a processor so that a computer implements the method for obtaining an image recognition model according to any one of claims 1-6 or the image recognition method according to any one of claims 7-9.

Citation Information

Patent Citations

  • Vocabulary entry classification method and audit information extraction method

    CN109635289A

  • Image recognition method and device, equipment and storage medium

    CN110175653A

  • Method and device for identifying image

    CN111582185A