Training method of image classification model, image classification method and device
By introducing textual feature descriptions into the training of image classification models and expanding the number of samples using pre-trained models, the semantic alignment ability between images and text is improved. This solves the problems of high cost and low accuracy of manual annotation in existing technologies and achieves higher fine-grained image classification accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2026-03-31
AI Technical Summary
Existing image classification models focus too much on image modalities during training, resulting in the need for a large amount of manually labeled data, which is costly and has low accuracy, affecting the accuracy of fine-grained image classification.
By introducing textual feature descriptions, the pre-trained model is trained based on multiple image-text pairs, expanding the number of samples, improving the semantic alignment between images and text, reducing the dependence on sample image labels, and using natural language descriptions to provide prior knowledge, thereby enhancing the model's ability to focus on features of subdivided categories.
It improves the accuracy of image classification models, reduces reliance on manual annotation, and enhances the accuracy of fine-grained image classification.
Smart Images

Figure CN116543264B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of image processing technology, and in particular to a training method for an image classification model, an image classification method, and an apparatus. Background Technology
[0002] Fine-grained image categorization, also known as sub-category recognition, has become a popular research area in computer vision and pattern recognition in recent years. Currently, the main approach to fine-grained image categorization is to use trained image classification models. However, because current model training focuses only on image modalities, such as the beak and limbs of a bird in an image, it requires a large amount of manually labeled data. This is not only costly but also inevitably introduces errors during the manual labeling process, potentially reducing the accuracy of the image classification model and consequently leading to lower accuracy in fine-grained image classification results. Therefore, improving the accuracy of the model, and thus the accuracy of fine-grained image classification results, is an urgent technical problem to be solved. Summary of the Invention
[0003] This specification provides one or more embodiments of a method for training an image classification model. The method includes acquiring n sample images and M texts from a training sample set. The training sample set includes N sample images and the M texts. The M texts correspond one-to-one with M subcategories of a target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, where n is less than N. The n sample images and M texts are input into a model to be trained for training processing to obtain an image classification result for each sample image. The model to be trained is constructed based on a pre-trained model. The pre-trained model is trained based on multiple image-text pairs. If a preset training termination condition is met based on the image classification results, an image classification model is obtained from the model to be trained. The image classification model is used to classify the input target image to be classified, obtaining the target subcategories to which the target image belongs under the target category.
[0004] This specification provides one or more embodiments of an image classification method. The method includes acquiring a target image to be classified; classifying the target image using a pre-trained image classification model to obtain a classification result, the classification result representing the target image's sub-category among multiple sub-categories included in the target category. The image classification model is obtained by training a model to be trained using n sample images and M texts obtained from a training sample set. The training sample set includes N sample images and the M texts. The M texts correspond one-to-one with the M sub-categories of the target category. The texts are used to describe the characteristics of the corresponding sub-categories. n, N, and M are integers greater than 1, where n is less than N. The model to be trained is constructed based on a pre-trained model, which is trained based on multiple image-text pairs.
[0005] This specification provides one or more embodiments of a training apparatus for an image classification model. The apparatus includes an acquisition module that acquires n sample images and M texts from a training sample set. The training sample set includes the N sample images and the M texts. The M texts correspond one-to-one with M subcategories of a target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The apparatus also includes a training module that inputs the n sample images and the M texts into a model to be trained for training processing to obtain an image classification result for each sample image. The model to be trained is constructed based on a pre-trained model. The pre-trained model is obtained by training multiple image-text pairs. The apparatus also includes an acquisition module that, if a preset training termination condition is met based on the image classification results, acquires the image classification model from the model to be trained.
[0006] This specification provides one or more embodiments of an image classification apparatus. The apparatus includes an acquisition module for acquiring a target image to be classified. The apparatus also includes a classification module for classifying the target image using a pre-trained image classification model to obtain a classification result. The classification result characterizes the target image's sub-category among multiple sub-categories included in the target category. The image classification model is obtained by training a model to be trained using n sample images and M texts obtained from a training sample set. The training sample set includes N sample images and the M texts. The M texts correspond one-to-one with the M sub-categories of the target category. The texts are used to describe the characteristics of the corresponding sub-categories. n, N, and M are integers greater than 1, where n is less than N. The model to be trained is constructed based on a pre-trained model, which is trained using multiple image-text pairs.
[0007] This specification provides one or more embodiments of a training apparatus for an image classification model. The apparatus includes a processor. The apparatus also includes a memory arranged to store computer-executable instructions. When executed, the computer-executable instructions cause the processor to acquire n sample images and M texts from a training sample set. The training sample set includes the N sample images and the M texts. The M texts correspond one-to-one with M subcategories of a target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, where n is less than N. The n sample images and the M texts are input into a model to be trained for training processing to obtain an image classification result for each sample image. The model to be trained is constructed based on a pre-trained model. The pre-trained model is trained based on multiple image-text pairs. If a preset training termination condition is met based on the image classification results, an image classification model is obtained from the model to be trained. The image classification model is used to classify the input target image to be classified, obtaining the target subcategories to which the target image belongs under the target category.
[0008] This specification provides one or more embodiments of a training apparatus for an image classification model. The apparatus includes a processor. It also includes a memory arranged to store computer-executable instructions. When executed, the computer-executable instructions cause the processor to acquire a target image to be classified. The target image is classified using a pre-trained image classification model to obtain a classification result, which characterizes the target image's sub-category among multiple sub-categories included in the target category. The image classification model is obtained by training a model to be trained using n sample images and M texts obtained from a training sample set. The training sample set includes N sample images and the M texts. The M texts correspond one-to-one with the M sub-categories of the target category. The texts are used to describe the characteristics of the corresponding sub-categories. n, N, and M are integers greater than 1, where n is less than N. The model to be trained is constructed based on a pre-trained model, which is trained using multiple image-text pairs.
[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions. When executed by a processor, the computer-executable instructions acquire n sample images and M texts from a training sample set. The training sample set includes N sample images and the M texts. The M texts correspond one-to-one with M subcategories of a target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, where n is less than N. The n sample images and the M texts are input into a training model for training processing to obtain an image classification result for each sample image. The training model is constructed based on a pre-trained model. The pre-trained model is trained based on multiple image-text pairs. If a preset training termination condition is met based on the image classification results, an image classification model is obtained from the training model. The image classification model is used to classify the input target image to obtain the target subcategories to which the target image belongs under the target category.
[0010] This specification provides one or more embodiments of a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When executed by a processor, the computer-executable instructions acquire a target image to be classified. The target image is classified using a pre-trained image classification model to obtain a classification result, which characterizes the target image's sub-category among multiple sub-categories included in the target category. The image classification model is obtained by training a model to be trained using n sample images and M texts obtained from a training sample set. The training sample set includes N sample images and the M texts. The M texts correspond one-to-one with the M sub-categories of the target category. The texts are used to describe the characteristics of the corresponding sub-categories. n, N, and M are integers greater than 1, where n is less than N. The model to be trained is constructed based on a pre-trained model, which is trained using multiple image-text pairs. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the first process of training an image classification model according to an embodiment of this specification;
[0013] Figure 2 This is a schematic diagram of a second process for training an image classification model provided in an embodiment of this specification;
[0014] Figure 3 This is a schematic diagram of the structure of an image classification model provided in an embodiment of this specification;
[0015] Figure 4 This is a schematic diagram of a third process for training an image classification model provided in an embodiment of this specification;
[0016] Figure 5 A flowchart illustrating an image classification method provided in an embodiment of this specification;
[0017] Figure 6 A schematic diagram of the module composition of a training device for an image classification model provided in an embodiment of this specification;
[0018] Figure 7 This is a schematic diagram of the module composition of an image classification device provided in the embodiments of this specification;
[0019] Figure 8 A schematic diagram of the structure of a training device for an image classification model provided in an embodiment of this specification;
[0020] Figure 9 This is a schematic diagram of the structure of an image classification device provided in an embodiment of this specification. Detailed Implementation
[0021] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0022] Figure 1 This is a schematic flowchart illustrating a training method for an image classification model provided in one or more embodiments of this specification. Figure 1 The method described above can be executed by a training device for the image classification model (hereinafter referred to as the training device), which can be located on a terminal device or a server. The terminal device can be a mobile phone, tablet, desktop computer, laptop, etc.; the server can be a standalone server or a server cluster consisting of multiple servers. For example... Figure 1 As shown, the method includes the following steps:
[0023] Step S102: Obtain n sample images and M texts from the training sample set; the training sample set includes N sample images and M texts, and the M texts correspond one-to-one with the M subcategories of the target category. Each text is used to describe the characteristics of the corresponding subcategory; n, N, and M are integers greater than 1, and n is less than N.
[0024] Specifically, N sample images and M texts are acquired, and a training sample set is constructed based on these N sample images and M texts. In each training step, n sample images and M texts are retrieved from the training sample set for training in the current training step, and the acquired n sample images and M texts are used for training in the current training step, i.e., subsequent steps S104 and S106 are executed. In other words, the M texts in the training sample set participate in the training process in each training step. A training step refers to the period from when training data is input into the model to when the model's parameters are adjusted; specifically, the last training step refers to the period from when training data is input into the model to when the training termination condition is met. Each sample image carries a label, which represents the category to which the sample image belongs. It can be understood that when the category represented by the label is a subcategory of the target category, the corresponding sample image is a positive sample; when the category represented by the label is not a subcategory of the target category, the corresponding sample image is a negative sample. Typically, M is less than N, and the values of n, N, and M can be set as needed in practical applications. This manual does not impose specific restrictions on these values.
[0025] Furthermore, acquiring M texts may include: for each of the M subcategories of the target category, retrieving statements and / or paragraphs from a specified location that describe the characteristics of that subcategory, and identifying the retrieved statements and / or paragraphs as the text corresponding to that subcategory. The target category can be set as needed in practical applications, such as birds, vehicles, plants, etc. The subcategories can be first-level subcategories of the target category, second-level subcategories further subdivided from the first-level subcategories, or even finer divisions, etc., which are not specifically limited in this specification. It is understood that the subcategories differ depending on the target category, and the descriptive content of the text may also differ depending on the subcategories. For example, if the target category is birds, and the subcategories can include ostriches, wingless birds, eagles, moa, etc., the text can be used to describe the appearance (e.g., eye color, beak length, feather color, etc.) and habitat of birds belonging to the corresponding subcategories; similarly, if the target category is vehicles, and the subcategories include cars, commercial vehicles, tractors, motorcycles, etc., the text can be used to describe the body, lights, wheels, etc. of vehicles belonging to the corresponding subcategories. The specified location can be the internet, documents, etc., and can be set as needed in practical applications.
[0026] Therefore, by introducing text to describe the characteristics of each sub-category, the natural language description can provide effective prior knowledge to the model to be trained, enabling the model to better focus on the essential features and characteristics related to each sub-category, thereby improving the model's classification accuracy and reducing the model's dependence on the labels of sample images.
[0027] Step S104: Input n sample images and M texts into the model to be trained for training processing to obtain the image classification result of each sample image; the model to be trained is built based on the pre-trained model, which is trained based on multiple image-text pairs;
[0028] Specifically, after obtaining n sample images and M texts from the training sample set, the obtained n sample images and M texts are input into the model to be trained for training at the current training step.
[0029] To address the high cost and low accuracy issues associated with extensive manual annotation, the training model in this embodiment is built upon a pre-trained model. This pre-trained model is trained on multiple image-text pairs, utilizing a large number (hundreds of millions) of readily available image-text pairs. In an unsupervised manner, pre-training tasks such as image-to-text and text-to-image generation are established. This significantly expands the sample size and enhances the robustness of the pre-trained model in aligning image and text semantics. This improves the image and text feature extraction capabilities of the training model built upon this pre-trained model and further reduces the training model's dependence on sample image labels.
[0030] Step S106: If the preset training termination condition is met based on the image classification result, the image classification model is obtained from the model to be trained. The image classification model is used to classify the input target image to be classified and obtain the target sub-category to which the target image belongs under the target category.
[0031] Specifically, using a determined loss function, loss calculation is performed based on the label and image classification result of each sample image to obtain the target loss. It is then determined whether the target loss is less than a preset loss. If so, the training termination condition is met, and an image classification model is obtained from the training model. If not, the next training step is performed, i.e., the process returns to step S102. In one implementation, the determined loss function can be ASL Loss (Asymmetric Loss). It should be noted that the loss function can be set according to needs in practical applications.
[0032] In one or more embodiments of this specification, n sample images and M texts are obtained from a training sample set. These n sample images and M texts are then input into the model to be trained for training processing to obtain the image classification result for each sample image. If the image classification result determines that a preset training termination condition is met, an image classification model is obtained from the model to be trained. The training sample set includes N sample images and M texts, where the M texts correspond one-to-one with the M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The model to be trained is constructed based on a pre-trained model, which is obtained by training multiple image-text pairs. Since the pre-trained model is obtained through unsupervised training on multiple image-text pairs, it greatly expands the number of samples and makes the pre-trained model more robust to the alignment of image and text semantics. This improves the ability of the model to be trained based on the pre-trained model to extract image and text features, ensuring accurate image classification and reducing the dependence of the model to be trained on sample image labels. By introducing text to describe the characteristics of each sub-category, the model can be provided with effective prior knowledge through natural language descriptions, allowing it to better focus on the essential characteristics related to each sub-category. This further reduces the dependence of the model to be trained on sample image labels and improves the accuracy of the model to be trained, thereby improving the accuracy of the resulting image classification model and ensuring accurate fine-grained image classification.
[0033] In one or more embodiments of the specification, during the training process, feature extraction is performed first, and then image classification is performed based on the extracted features. Specifically, as shown... Figure 2 As shown, step S104 may include steps S104-2 and S104-4:
[0034] Step S104-2: Input n sample images and M texts into the model to be trained, and use the model to extract features from the n sample images and M texts to obtain the target image features of the n sample images and the target text features of the M texts.
[0035] like Figure 3 As shown, the model to be trained includes an image processing module and a text processing module, and correspondingly, as... Figure 4 As shown, step S104-2 may include steps S104-22 and S104-24:
[0036] Step S104-22: Input n sample images into the image processing module of the model to be trained, and perform image feature extraction processing on the n sample images through the image processing module to obtain the target image features of the n sample images;
[0037] Specifically, such as Figure 3 As shown, the pre-trained model includes an image encoder, and the image processing module includes a visual encoder and the image encoder in the pre-trained model. Accordingly, steps S104-22 may include: transforming n sample images through the image encoder to obtain a first image feature; transforming n sample images through the visual encoder to obtain a second image feature; and fusing the first image feature and the second image feature to obtain the target image feature of the n sample images.
[0038] The process of transforming n sample images by the image encoder is the same as that of the visual encoder, except that the parameters used for the transformation are different. Since the pre-trained model is pre-trained, it can be understood that during the training of the model to be trained, the parameters of the image encoder in the pre-trained model remain fixed, while the parameters of the visual encoder are continuously adjusted and optimized as the number of training steps increases.
[0039] Furthermore, the transformation processing of the n sample images may include: segmenting each sample image to obtain T sample sub-images, where T is an integer greater than 1; multiplying each sample sub-image with a preset transformation vector to obtain the sub-image features of each sample sub-image; and generating corresponding image features based on the sub-image features. It can be understood that when the image encoder performs transformation processing on the n sample images, the aforementioned generation of corresponding image features based on the sub-image features constitutes the generation of the first image features; when the visual encoder performs transformation processing on the n sample images, the aforementioned generation of corresponding image features based on the sub-image features constitutes the generation of the second image features.
[0040] In one implementation, segmenting each sample image may include segmenting the length and width of each sample image. The first image feature and the second image feature may be in matrix form. Fusion processing of the first image feature and the second image feature may include adding elements at the same positions in the first image feature and the second image feature.
[0041] As an example, n is 10; the size of each sample image is 16*16*3, where the first 16 represents the length of the sample image, the second 16 represents the width of the sample image, and 3 represents the number of channels; the segmentation process involves dividing the length and width of each sample image into two equal parts; therefore, for the image encoder, segmenting each sample image results in four sample sub-images of size 8*8*3 for each sample image. The sample sub-images of the 10 sample images can be represented as 10*4*(8*8*3), i.e., 10*4 Each sample sub-image can be represented as 1*192. Multiplying each sample sub-image by a preset transformation vector yields a sub-image feature of 1*1024. The first image feature generated based on these features is 10*4*1024, meaning it represents the features of each sub-image corresponding to the 10 sample images. Similarly, each sub-image feature of each sample image can be represented as 4*1024. Likewise, the second image feature obtained by the visual encoder can be represented as 10*4*1024. Fusing the first and second image features yields a target image feature of 10*4*1024. It is understood that the values of elements at the same positions in the first and second image features may be the same or different.
[0042] Steps S104-24: Input the M texts into the text processing module of the model to be trained, and perform text feature extraction on the M texts to obtain the target text features of the M texts.
[0043] Specifically, such as Figure 3 As shown, the pre-trained model also includes a text encoder, and the text processing module includes a transformer and the text encoder in the pre-trained model. Accordingly, steps S104-24 may include: encoding each of the M texts using the text encoder to obtain a first text feature; encoding each of the M texts using the transformer to obtain a second text feature; and fusing the first text feature and the second text feature to obtain the target text feature of the M texts.
[0044] More specifically, a text encoder encodes each of the M texts to obtain a first sub-text feature for each text; based on the M first sub-text features, a first text feature is generated; a converter encodes each of the M texts to obtain a second sub-text feature for each text; based on the M second sub-text features, a second text feature is generated; and the sub-text features at the same positions of the first and second text features are fused to obtain the target text features for the M texts. The fusion result of the sub-text features at the same positions of the first and second text features is called the text fusion feature; that is, the target text features include M text fusion features.
[0045] The process of encoding each text element by the text encoder is the same as that by the converter, except that the parameters used for encoding differ. Since the pre-trained model is pre-trained, it can be understood that during the training of the model to be trained, the parameters of the text encoder in the pre-trained model remain fixed, while the parameters of the converter are continuously adjusted and optimized as the number of training steps increases. The specific encoding process can be set as needed in practical applications; this specification does not impose specific limitations on it.
[0046] In one implementation, the first text feature and the second text feature can be in matrix form. The fusion processing of the first text feature and the second text feature can include adding the values of elements at the same position in the first text feature and the second text feature.
[0047] Continuing with the example above, with M=5, each text is encoded by a text encoder to obtain a first sub-text feature of 1*1024 for each text. The first text feature generated based on the 5 first sub-text features can be represented as 5*1024. Each text is then encoded by a converter to obtain a second sub-text feature of 1*1024 for each text. The second text feature generated based on the 5 second sub-text features can be represented as 5*1024. The first and second text features are then fused to obtain the target text feature, which can be represented as 5*1024.
[0048] Step S104-4: Using the model to be trained, perform image classification processing based on the target image features and target text features to obtain the classification result for each sample image.
[0049] like Figure 3 As shown, the model to be trained includes a fusion module, and correspondingly, as... Figure 4 As shown, step S104-4 may include steps S104-42 to S104-46:
[0050] Step S104-42: The target image features and target text features are concatenated to obtain concatenated features;
[0051] Specifically, the sub-image features of each sample image in the target image features are concatenated with the target text features to obtain sub-concatenated features; and concatenated features are generated based on the sub-concatenated features of each sample image.
[0052] Continuing with the example above, the target image features are 10*4*1024, where each sub-image feature of each sample image is 4*1024. The sub-image features of each sample image, 4*1024, are concatenated with the target text features, 5*1024, to obtain a sub-concatenation feature pair of 9*1024. The generated concatenation feature is 10*9*1024; that is, the concatenation feature is a representation of the sub-concatenation features of each sample image.
[0053] Steps S104-44: Through the fusion module of the model to be trained, correlation learning processing is performed based on the spliced features to obtain the correlation learning results;
[0054] As described above, the splicing features include T sub-image features corresponding to each sample image and M text fusion features corresponding to each sample image. Accordingly, steps S104-44 may include:
[0055] The T sub-image features and M text fusion features corresponding to each sample image in the spliced features are determined as the features to be learned for each sample image. Each feature to be learned is then determined as the target feature to be learned. The target feature to be learned is then multiplied by each feature to be learned in the corresponding sample image to obtain the first dot product result. Each first dot product result is then normalized to obtain the first normalized result. The first normalized result is then determined as the weight of the corresponding feature to be learned. The corresponding features to be learned are then weighted to obtain the first sub-correlation learning result of the target feature to be learned. Based on each first sub-correlation learning result, the correlation learning result is generated.
[0056] The process of generating relevance learning results based on the first sub-relevance learning results may include: determining the second sub-relevance learning result for each sample image based on the first sub-relevance learning results; generating the third sub-relevance learning result based on the second sub-relevance learning results for each sample image; and determining the learning results corresponding to M texts in the third sub-relevance learning results as the relevance learning results.
[0057] In the embodiments of this specification, for the convenience of data processing, each first normalization result obtained by normalization processing belongs to 0 to 1.
[0058] Continuing the example above, n = 10, T = 4, M = 5; for ease of description, each sample image is denoted as sample image 1, sample image 2...sample image 10, and the four sub-image features of each sample image are denoted as sub-image feature 1, sub-image feature 2, sub-image feature 3, and sub-image feature 4, and the M text fusion features are denoted as text fusion feature 1, text fusion feature 2...text fusion feature 5. Taking sample image 1 as an example, firstly, sub-image feature 1 of sample image 1 is determined as the target feature to be learned. Then, sub-image feature 1 of sample image 1 is subjected to dot product processing with sub-image feature 1, sub-image feature 2, sub-image feature 3, sub-image feature 4, text fusion feature 1, text fusion feature 2...text fusion feature 5 of sample image 1, respectively, to obtain the corresponding 9 first dot product processing results; the 9 first dot product processing results are then normalized to obtain 9 first normalized results; for example, following the aforementioned dot product processing order, the 9 first dot product processing results are 0.1, 0... Given the values 0.3, 0.2, 0.5, 0.7, 0.4, 0.9, 0.1, and 0.4, we assign 0.1 as the weight of sub-image feature 1, 0.3 as the weight of sub-image feature 2, and so on, assigning 0.4 as the weight of text fusion feature 5. We then perform weighted processing on the four sub-image features and five text fusion features corresponding to sample image 1 to obtain the first sub-relevance learning result of sub-image feature 1 for sample image 1. That is, the first sub-relevance learning result of sub-image feature 1 for sample image 1 = 0.1 * sub-image feature 1 + 0.3 * sub-image feature 2 + ... + 0.4 * text fusion feature 5. Similarly, we can obtain the first sub-relevance learning result for each feature to be learned. Each first sub-correlation learning result can be represented as 1*1024; the second sub-correlation learning result can be represented as 9*1024; the third sub-correlation learning result can be represented as 10*9*1024; and the correlation learning result can be represented as 10*5*1024, where 10 represents 10 samples and 5 represents 5 texts, i.e. 5 subcategories.
[0059] Steps S104-46: Based on the correlation learning results, determine the classification result for each sample image.
[0060] Specifically, for each sample image, the target relevance learning result for each sub-category is obtained from the relevance learning results; each target relevance learning result is then multiplied by a preset classification vector to obtain a second dot product result; each second dot product result is normalized to obtain a second normalized result; the second normalized result is determined as the probability that the sample image belongs to the corresponding sub-category; and based on the probabilities corresponding to the sample image, a classification result for the sample image is generated. The second normalized result is between 0 and 1; the classification result of the sample image can be a one-dimensional vector composed of the probabilities corresponding to the sample image.
[0061] Continuing with the example above, the relevance learning result is represented as 10*5*1024. Taking sample image 1 as an example, the target relevance learning result of sample image 1 for each of the 5 subcategories can be represented as 1*1024. Each target relevance learning result is then multiplied by a preset classification vector to obtain 5 second dot product results. Each second dot product result is then normalized to obtain 5 second normalized results. These 5 second normalized results are used to determine the probability that sample image 1 belongs to one of the 5 subcategories. The classification result of sample image 1 can be represented as a one-dimensional vector consisting of 5 probabilities, i.e., 1*5. Similarly, the image classification result for each sample image can be obtained.
[0062] Therefore, the image classification result for each sample image is determined based on the correlation learning mechanism, ensuring the accuracy of the image classification results.
[0063] Since the text processing module of the model under training has learned the text features of each subcategory under the target category by the end of model training, that is, it has learned the accurate target text features, in order to improve the efficiency of subsequent image classification processing based on the obtained image classification model without extracting the target text features from the M texts again, in one or more embodiments of this specification, obtaining the image classification model from the model under training in step S106 includes: associating the current target text features with the current image processing module and the current fusion module to obtain the image classification model. That is, the image classification model includes the aforementioned image processing module and fusion module. The target text features can be stored in the image classification model or in other storage space, as long as they are associated with the image classification model so that the image classification model can obtain the target text features when performing image classification processing.
[0064] In one or more embodiments of this specification, n sample images and M texts are obtained from a training sample set. These n sample images and M texts are then input into the model to be trained for training processing to obtain the image classification result for each sample image. If the image classification result determines that a preset training termination condition is met, an image classification model is obtained from the model to be trained. The training sample set includes N sample images and M texts, where the M texts correspond one-to-one with the M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The model to be trained is constructed based on a pre-trained model, which is obtained by training multiple image-text pairs. Since the pre-trained model is obtained through unsupervised training on multiple image-text pairs, it greatly expands the number of samples and makes the pre-trained model more robust to the alignment of image and text semantics. This improves the ability of the model to be trained based on the pre-trained model to extract image and text features, ensuring accurate image classification and reducing the dependence of the model to be trained on sample image labels. By introducing text to describe the characteristics of each sub-category, the model can be provided with effective prior knowledge through natural language descriptions, allowing it to better focus on the essential characteristics related to each sub-category. This further reduces the dependence of the model to be trained on sample image labels and improves the accuracy of the model to be trained, thereby improving the accuracy of the resulting image classification model and ensuring accurate fine-grained image classification.
[0065] Based on the same technical concept, one or more embodiments of this specification also provide an image classification method corresponding to the training method of the image classification model described above. Figure 5 This is a flowchart illustrating an image classification method provided in one or more embodiments of this specification. Figure 5 The method described can be executed by an image classification device; this device can be located in a terminal device or a server. The terminal device can be a mobile phone, tablet, desktop computer, laptop, etc.; the server can be a standalone server or a server cluster consisting of multiple servers. Figure 5 As shown, the method includes the following steps:
[0066] Step S202: Obtain the target image to be classified;
[0067] The target image can be any image, and may include objects belonging to a subcategory of the target category (such as birds, vehicles, etc.), or may not include objects belonging to a subcategory of the target category.
[0068] Step S204: Classify the target image using a pre-trained image classification model to obtain the classification result of the target image; the classification result represents the target sub-category to which the target image belongs among the multiple sub-categories included in the target category.
[0069] The image classification model is obtained by training the model using n sample images and M texts obtained from the training sample set. The training sample set includes N sample images and M texts, with each of the M texts corresponding to one of the M subcategories of the target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The model to be trained is built based on the pre-trained model, which is trained on multiple image-text pairs.
[0070] In one or more embodiments of this specification, the image classification model includes an image processing module and a fusion module; correspondingly, step S204 may include:
[0071] The target image is input into the image processing module, which performs image feature extraction processing on the target image to obtain the target image features.
[0072] The target image features and the acquired target text features are concatenated to obtain concatenated features.
[0073] The correlation learning results are obtained by using the fusion module to perform correlation learning processing based on the spliced features.
[0074] Based on the relevance learning results, the classification result of the target image is determined.
[0075] Furthermore, the classification result of the target image determined based on the relevance learning results can include:
[0076] Obtain the target relevance learning results of the target image with respect to each sub-category from the relevance learning results;
[0077] The dot product of each target relevance learning result and the preset classification vector is performed to obtain the third dot product result.
[0078] Normalize each third dot product result to obtain the third normalized result;
[0079] The third normalization result is determined as the probability that the target image belongs to the corresponding sub-category;
[0080] The sub-category corresponding to the highest probability is determined as the classification result of the target image; or, the sub-category corresponding to the target probability greater than a preset probability threshold is determined as the classification result of the target image.
[0081] As mentioned earlier, each probability is between 0 and 1. It can be understood that when the probability is 0, it indicates that the target image does not belong to the corresponding sub-category. The larger the probability value, the greater the probability that the target image belongs to the corresponding sub-category. When all probabilities are 0, it indicates that the target image does not belong to the target category. The corresponding classification result of the target image indicates that the target image does not belong to any sub-category under the target category.
[0082] It should be noted that the specific implementation process of each of the above steps can be found in the relevant descriptions above, and the repeated parts will not be repeated here.
[0083] In one or more embodiments of this specification, a target image to be classified is obtained; the target image is classified using a pre-trained image classification model to obtain a classification result; the classification result represents the target sub-category to which the target image belongs among multiple sub-categories included in the target category. Since the image classification model used for classification is trained on a model to be trained based on a pre-trained model, and since the pre-trained model is unsupervised training based on multiple image-text pairs, it greatly expands the number of samples, making the pre-trained model more robust to the alignment of image semantics and text semantics. This improves the ability of the model to extract image and text features based on the pre-trained model, ensuring accurate image classification. Furthermore, during training, text is introduced to describe the characteristics of each sub-category, providing effective prior knowledge to the model to be trained through natural language descriptions. This allows the model to better focus on the essential characteristics related to each sub-category, further reducing the model's dependence on the labels of sample images and improving the accuracy of the model. Therefore, image classification is performed based on this highly accurate image classification model, achieving accurate, fine-grained image classification.
[0084] Based on the same technical concept, one or more embodiments of this specification also provide a training device for an image classification model, corresponding to the training method of the image classification model described above. Figure 6 This is a schematic diagram illustrating the module composition of a training device for an image classification model provided in one or more embodiments of this specification, as shown below. Figure 6 As shown, the device includes:
[0085] The acquisition module 301 acquires n sample images and M texts from the training sample set; the training sample set includes N sample images and the M texts, the M texts correspond one-to-one with the M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories; n, N and M are integers greater than 1, and n is less than N;
[0086] Training module 302 inputs the n sample images and the M texts into the model to be trained for training processing to obtain the image classification result of each sample image; the model to be trained is constructed based on the pre-trained model, which is trained based on multiple image-text pairs;
[0087] If the image classification result determines that the preset training termination condition is met, the image classification model is obtained from the model to be trained by the acquisition module 303. The image classification model is used to classify the input target image to be classified and obtain the target sub-category to which the target image belongs under the target category.
[0088] The image classification model training apparatus provided in one or more embodiments of this specification acquires n sample images and M texts from a training sample set, inputs the n sample images and M texts into the model to be trained for training processing, and obtains the image classification result for each sample image; if the image classification result determines that a preset training termination condition is met, the image classification model is obtained from the model to be trained. The training sample set includes N sample images and M texts, where the M texts correspond one-to-one with the M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories; n, N, and M are integers greater than 1, and n is less than N; the model to be trained is constructed based on a pre-trained model, which is obtained by training multiple image-text pairs. Since the pre-trained model is obtained through unsupervised training on multiple image-text pairs, it greatly expands the number of samples and makes the pre-trained model more robust to the alignment of image and text semantics. This improves the ability of the model to be trained based on the pre-trained model to extract image and text features, ensuring accurate image classification and reducing the dependence of the model to be trained on sample image labels. By introducing text to describe the characteristics of each sub-category, the model can be provided with effective prior knowledge through natural language descriptions, allowing it to better focus on the essential characteristics related to each sub-category. This further reduces the dependence of the model to be trained on sample image labels and improves the accuracy of the model to be trained, thereby improving the accuracy of the resulting image classification model and ensuring accurate fine-grained image classification.
[0089] It should be noted that the embodiments of the training device for the image classification model in this specification and the embodiments of the training method for the image classification model in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding image classification model training method mentioned above, and the repeated parts will not be described again.
[0090] Furthermore, corresponding to the image classification method described above, based on the same technical concept, one or more embodiments of this specification also provide an image classification device. Figure 7 This specification provides a schematic diagram of the module composition of an image classification device according to one or more embodiments. Figure 7 As shown, the device includes:
[0091] Module 401 acquires the target image to be classified;
[0092] The classification module 402 performs classification processing on the target image using a pre-trained image classification model to obtain the classification result of the target image; the classification result represents the target sub-category to which the target image belongs among multiple sub-categories included in the target category;
[0093] The image classification model is obtained by training the model to be trained using n sample images and M texts obtained from the training sample set. The training sample set includes N sample images and the M texts, and the M texts correspond one-to-one with the M subcategories of the target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The model to be trained is built based on a pre-trained model, which is trained based on multiple image-text pairs.
[0094] The image classification apparatus provided in one or more embodiments of this specification acquires a target image to be classified; classifies the target image using a pre-trained image classification model to obtain a classification result; the classification result represents the target sub-category to which the target image belongs among multiple sub-categories included in the target category. Since the image classification model used in the classification process is trained on a model to be trained based on a pre-trained model, and since the pre-trained model is unsupervised training based on multiple image-text pairs, it greatly expands the number of samples, making the pre-trained model more robust to the alignment of image semantics and text semantics. This improves the ability of the model to extract image and text features based on the pre-trained model, ensuring accurate image classification. Furthermore, during training, text is introduced to describe the characteristics of each sub-category, providing effective prior knowledge to the model to be trained through natural language descriptions. This allows the model to better focus on the essential characteristics related to each sub-category, further reducing the model's dependence on the labels of sample images and improving the accuracy of the model. Therefore, based on this highly accurate image classification model, accurate fine-grained image classification is achieved.
[0095] It should be noted that the embodiments of the image classification device in this specification and the embodiments of the image classification method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding image classification method mentioned above, and the repeated parts will not be described again.
[0096] Furthermore, corresponding to the image classification model training method described above, based on the same technical concept, one or more embodiments of this specification also provide an image classification model training device, which is used to execute the above-described image classification model training method. Figure 8 This is a schematic diagram of the structure of a training device for an image classification model provided in one or more embodiments of this specification.
[0097] like Figure 8 As shown, the training device for the image classification model can vary significantly due to differences in configuration or performance. It may include one or more processors 501 and memory 502, with memory 502 storing one or more application programs or data. Memory 502 can be temporary or persistent storage. The application programs stored in memory 502 may include one or more modules (not shown), each module including a series of computer-executable instructions for the image classification model training device. Furthermore, processor 501 may be configured to communicate with memory 502, executing the series of computer-executable instructions in memory 502 on the image classification model training device. The image classification model training device may also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input / output interfaces 505, one or more keyboards 506, etc.
[0098] In one specific embodiment, the training device for the image classification model includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the training device for the image classification model, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:
[0099] Obtain n sample images and M texts from the training sample set; the training sample set includes N sample images and M texts, the M texts correspond one-to-one with M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories; n, N and M are integers greater than 1, and n is less than N;
[0100] The n sample images and M texts are input into the training model for training processing to obtain the image classification result of each sample image; the training model is built based on the pre-trained model, which is trained based on multiple image-text pairs;
[0101] If the preset training termination condition is met based on the image classification result, then the image classification model is obtained from the model to be trained; the image classification model is used to classify the input target image to be classified, and to obtain the target sub-category to which the target image belongs under the target category.
[0102] The image classification model training device provided in one or more embodiments of this specification acquires n sample images and M texts from a training sample set, inputs the n sample images and M texts into the model to be trained for training processing, and obtains the image classification result for each sample image; if the image classification result determines that a preset training termination condition is met, the image classification model is obtained from the model to be trained. The training sample set includes N sample images and M texts, where the M texts correspond one-to-one with the M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories; n, N, and M are integers greater than 1, and n is less than N; the model to be trained is constructed based on a pre-trained model, which is obtained by training multiple image-text pairs. Since the pre-trained model is obtained through unsupervised training on multiple image-text pairs, it greatly expands the number of samples and makes the pre-trained model more robust to the alignment of image and text semantics. This improves the ability of the model to be trained based on the pre-trained model to extract image and text features, ensuring accurate image classification and reducing the dependence of the model to be trained on sample image labels. By introducing text to describe the characteristics of each sub-category, the model can be provided with effective prior knowledge through natural language descriptions, allowing it to better focus on the essential characteristics related to each sub-category. This further reduces the dependence of the model to be trained on sample image labels and improves the accuracy of the model to be trained, thereby improving the accuracy of the resulting image classification model and ensuring accurate fine-grained image classification.
[0103] Furthermore, corresponding to the image classification method described above, based on the same technical concept, one or more embodiments of this specification also provide an image classification device for performing the above-described image classification method. Figure 9 This is a schematic diagram of the structure of an image classification device provided in one or more embodiments of this specification.
[0104] like Figure 9As shown, image classification devices can vary significantly due to differences in configuration or performance. They may include one or more processors 601 and memory 602, with memory 602 storing one or more application programs or data. Memory 602 can be temporary or persistent storage. The application programs stored in memory 602 may include one or more modules (not shown), each module including a series of computer-executable instructions for the image classification device. Furthermore, processor 601 may be configured to communicate with memory 602, executing the series of computer-executable instructions stored in memory 602 on the image classification device. The image classification device may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input / output interfaces 605, one or more keyboards 606, etc.
[0105] In one specific embodiment, the image classification device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the image classification device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:
[0106] Obtain the target image to be classified;
[0107] The target image is classified using a pre-trained image classification model to obtain a classification result for the target image; the classification result represents the target sub-category to which the target image belongs among multiple sub-categories included in the target category;
[0108] The image classification model is obtained by training the model to be trained using n sample images and M texts obtained from the training sample set. The training sample set includes N sample images and the M texts, and the M texts correspond one-to-one with the M subcategories of the target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The model to be trained is built based on a pre-trained model, which is trained based on multiple image-text pairs.
[0109] The image classification device provided in one or more embodiments of this specification acquires a target image to be classified; classifies the target image using a pre-trained image classification model to obtain a classification result; the classification result represents the target sub-category to which the target image belongs among multiple sub-categories included in the target category. Since the image classification model used for classification is trained on a model to be trained based on a pre-trained model, and since the pre-trained model is unsupervised training based on multiple image-text pairs, it greatly expands the number of samples, making the pre-trained model more robust to the alignment of image semantics and text semantics. This improves the ability of the model to extract image and text features based on the pre-trained model, ensuring accurate image classification. Furthermore, during training, text is introduced to describe the characteristics of each sub-category, providing effective prior knowledge to the model to be trained through natural language descriptions. This allows the model to better focus on the essential characteristics related to each sub-category, further reducing the model's dependence on the labels of sample images and improving the accuracy of the model. Therefore, based on this highly accurate image classification model, accurate fine-grained image classification is achieved.
[0110] It should be noted that the embodiments of the training device and image classification device for the image classification model described in this specification are based on the same inventive concept as the embodiments of the training method and image classification method for the image classification model described in this specification. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding training method and image classification method for the image classification model mentioned above, and the repeated parts will not be described again.
[0111] Furthermore, corresponding to the training method and image classification method of the image classification model described above, based on the same technical concept, one or more embodiments of this specification also provide a storage medium for storing computer-executable instructions. In a specific embodiment, the storage medium can be a USB flash drive, optical disk, hard disk, etc. When the computer-executable instructions stored in the storage medium are executed by the processor, they can realize the following process:
[0112] Obtain n sample images and M texts from the training sample set; the training sample set includes N sample images and M texts, the M texts correspond one-to-one with M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories; n, N and M are integers greater than 1, and n is less than N;
[0113] The n sample images and M texts are input into the training model for training processing to obtain the image classification result of each sample image; the training model is built based on the pre-trained model, which is trained based on multiple image-text pairs;
[0114] If the preset training termination condition is met based on the image classification result, then the image classification model is obtained from the model to be trained; the image classification model is used to classify the input target image to be classified, and to obtain the target sub-category to which the target image belongs under the target category.
[0115] The computer-executable instructions stored in the storage medium provided in one or more embodiments of this specification, when executed by a processor, acquire n sample images and M texts from a training sample set, input the n sample images and M texts into the model to be trained for training processing, and obtain the image classification result for each sample image; if it is determined from the image classification result that a preset training termination condition is met, then the image classification model is obtained from the model to be trained. The training sample set includes N sample images and M texts, where the M texts correspond one-to-one with the M subcategories of the target category, and the texts are used to describe the characteristics of the corresponding subcategories; n, N, and M are integers greater than 1, and n is less than N; the model to be trained is constructed based on a pre-trained model, which is obtained by training multiple image-text pairs. Since the pre-trained model is obtained through unsupervised training on multiple image-text pairs, it greatly expands the number of samples and makes the pre-trained model more robust to the alignment of image and text semantics. This improves the ability of the model to be trained based on the pre-trained model to extract image and text features, ensuring accurate image classification and reducing the dependence of the model to be trained on sample image labels. By introducing text to describe the characteristics of each sub-category, the model can be provided with effective prior knowledge through natural language descriptions, allowing it to better focus on the essential characteristics related to each sub-category. This further reduces the dependence of the model to be trained on sample image labels and improves the accuracy of the model to be trained, thereby improving the accuracy of the resulting image classification model and ensuring accurate fine-grained image classification.
[0116] In another specific embodiment, the storage medium can be a USB flash drive, optical disc, hard disk, etc., and the computer-executable instructions stored in the storage medium, when executed by a processor, can achieve the following process:
[0117] Obtain the target image to be classified;
[0118] The target image is classified using a pre-trained image classification model to obtain a classification result for the target image; the classification result represents the target sub-category to which the target image belongs among multiple sub-categories included in the target category;
[0119] The image classification model is obtained by training the model to be trained using n sample images and M texts obtained from the training sample set. The training sample set includes N sample images and the M texts, and the M texts correspond one-to-one with the M subcategories of the target category. The texts are used to describe the characteristics of the corresponding subcategories. n, N, and M are integers greater than 1, and n is less than N. The model to be trained is built based on a pre-trained model, which is trained based on multiple image-text pairs.
[0120] The computer-executable instructions stored in the storage medium provided in one or more embodiments of this specification, when executed by a processor, acquire a target image to be classified; classify the target image using a pre-trained image classification model to obtain a classification result of the target image; the classification result characterizes the target sub-category to which the target image belongs among multiple sub-categories included in the target category. The image classification model used in the classification process is trained on a pre-trained model. Since the pre-trained model is trained unsupervised on multiple image-text pairs, it significantly expands the sample size and enhances the robustness of the pre-trained model in aligning image and text semantics. This improves the image and text feature extraction capabilities of the model to be trained based on the pre-trained model, ensuring accurate image classification. Furthermore, during training, text is introduced to describe the characteristics of each sub-category, providing effective prior knowledge to the model through natural language descriptions. This allows the model to better focus on the essential characteristics of each sub-category, further reducing its dependence on the labels of sample images and improving its accuracy. Therefore, based on this highly accurate image classification model, accurate, fine-grained image classification is achieved.
[0121] It should be noted that the embodiments concerning storage media in this specification and the embodiments concerning the training method of image classification models in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding image classification model training method described above, and the repeated parts will not be described again.
[0122] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0123] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using a hardware physical module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0124] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0125] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0126] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0127] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0128] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0129] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0130] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0131] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0132] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0133] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0134] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0135] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0136] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0137] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. A method for training an image classification model, comprising: obtaining n sample images and M texts from a training sample set; the training sample set comprises N sample images and the M texts, the M texts correspond to M sub-categories of a target category one by one, and the texts are used for describing the characteristics of the corresponding sub-categories; n, N and M are integers greater than 1, and n is less than N; inputting the n sample images and the M texts into a to-be-trained model for training processing to obtain an image classification result of each sample image; the to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is obtained based on training of multiple image-text pairs; if it is determined that a preset training end condition is met according to the image classification result, an image classification model is obtained from the to-be-trained model; the image classification model is used for classifying a target image to be classified to obtain a target sub-category to which the target image belongs under the target category; wherein the to-be-trained model comprises a fusion module, the fusion module is used for performing correlation learning processing based on a spliced feature to obtain a correlation learning result, and the classification result of each sample image is determined according to the correlation learning result, and the spliced feature is a feature obtained by splicing a target image feature of the n sample images and a target text feature of the M texts; the determination of the classification result of each sample image according to the correlation learning result comprises: for each sample image, a target correlation learning result of the sample image with respect to each sub-category is obtained from the correlation learning result; performing dot product processing on each target correlation learning result and a preset classification vector to obtain a second dot product processing result; determining the probability that the sample image belongs to the corresponding sub-category according to the second dot product processing result, and determining the classification result of the sample image according to the probabilities corresponding to the sample image.
2. The method of claim 1, wherein the n sample images and the M texts are input into the to-be-trained model for training processing to obtain an image classification result of each sample image, comprising: inputting the n sample images and the M texts into the to-be-trained model, performing feature extraction processing on the n sample images and the M texts by the to-be-trained model to obtain target image features of the n sample images and target text features of the M texts; performing image classification processing based on the target image features and the target text features by the to-be-trained model to obtain an image classification result of each sample image.
3. The method of claim 2, wherein the to-be-trained model comprises an image processing module and a text processing module, and the feature extraction processing on the n sample images and the M texts by the to-be-trained model to obtain the target image features of the n sample images and the target text features of the M texts comprises: inputting the n sample images into the image processing module, performing image feature extraction processing on the n sample images by the image processing module to obtain target image features of the n sample images; inputting the M texts into the text processing module, performing text feature extraction processing on the M texts by the text processing module to obtain target text features of the M texts.
4. The method of claim 3, wherein the pre-trained model comprises an image encoder, and the image processing module comprises the image encoder and a visual encoder. The image feature extraction processing on the n sample images by the image processing module comprises: transforming the n sample images by the image encoder to obtain first image features; transforming the n sample images by the visual encoder to obtain second image features; fusing the first image features and the second image features to obtain target image features of the n sample images.
5. The method of claim 4, wherein the transforming the n sample images comprises: segmenting each of the sample images to obtain T sample sub-images of each of the sample images; T is an integer greater than 1; multiplying each of the sample sub-images with a preset transformation vector to obtain sub-image features of each of the sample sub-images; generating corresponding image features according to the sub-image features.
6. The method of claim 3, wherein the pre-trained model comprises a text encoder, and the text processing module comprises the text encoder and a converter; the text feature extraction processing on the M texts by the text processing module comprises: encoding each of the M texts by the text encoder to obtain first text features; encoding each of the M texts by the converter to obtain second text features; fusing the first text features and the second text features to obtain target text features of the M texts.
7. The method of claim 2, wherein the image classification processing based on the target image features and the target text features by the to-be-trained model to obtain a classification result of each of the sample images comprises: splicing the target image features and the target text features to obtain spliced features; performing correlation learning processing on the spliced features by the fusion module to obtain a correlation learning result; determining the classification result of each of the sample images according to the correlation learning result.
8. The method of claim 7, wherein the spliced features comprise T sub-image features corresponding to each of the sample images and M text fusion features corresponding to each of the sample images; and the correlation learning processing on the spliced features to obtain a correlation learning result comprises: determine the T sub-image features and the M text fusion features corresponding to each of the sample images in the splicing feature as the to-be-learned feature of each sample image; determine each of the to-be-learned features as a target to-be-learned feature in turn, perform dot product processing on the target to-be-learned feature and each to-be-learned feature of the corresponding sample image respectively to obtain a first dot product processing result; perform normalization processing on each of the first dot product processing results to obtain a first normalization result of each of the first dot product processing results; determine the first normalization result as the weight of the corresponding to-be-learned feature, and perform weighted processing on the corresponding to-be-learned features to obtain a first sub-correlation learning result of the target to-be-learned feature; generate a correlation learning result according to each of the first sub-correlation learning results.
9. The method of claim 7, wherein determining the probability that the sample image belongs to the corresponding fine classification category according to the second dot product processing result comprises: performing normalization processing on each of the second dot product processing results to obtain a second normalization result; determining the second normalization result as the probability that the sample image belongs to the corresponding fine classification category.
10. The method of claim 3, wherein the to-be-trained model further comprises a fusion module; and wherein obtaining an image classification model from the to-be-trained model comprises: associating the current target text feature with the current image processing module and the current fusion module to obtain an image classification model.
11. An image classification method, comprising: obtaining a target image to be classified; performing classification processing on the target image by using a pre-trained image classification model to obtain a classification result of the target image; the classification result representing a target fine classification category to which the target image belongs among a plurality of fine classification categories included in a target category; wherein the image classification model is obtained by training a to-be-trained model using n sample images and M texts obtained from a training sample set; the training sample set comprises N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of the target category, and the texts are used to describe the characteristics of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; the to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is trained based on a plurality of image-text pairs; wherein the image classification model comprises a fusion module, the fusion module is used for performing correlation learning processing based on splicing features to obtain a correlation learning result, and the classification result of the target image is determined according to the correlation learning result; the splicing features are features obtained by splicing a target image feature of the target image and a target text feature obtained; determining the classification result of the target image according to the correlation learning result comprises: obtaining a target correlation learning result of the target image with respect to each of the fine classification categories from the correlation learning result; performing dot product processing on each of the target correlation learning results and a preset classification vector to obtain a third dot product processing result; According to the third dot product processing result, a probability that the target image belongs to a corresponding fine classification category is determined, and a classification result of the target image is determined according to the probability that the target image belongs to the corresponding fine classification category.
12. The method of claim 11, wherein the image classification model comprises an image processing module; and the classification processing of the target image by the pre-trained image classification model comprises: inputting the target image into the image processing module, performing image feature extraction processing on the target image by the image processing module to obtain target image features of the target image; performing splicing processing on the target image features and the obtained target text features to obtain spliced features; performing correlation learning processing on the spliced features by the fusion module to obtain a correlation learning result; and determining the classification result of the target image according to the correlation learning result.
13. The method of claim 12, wherein the determining, according to the third dot product processing result, of a probability that the target image belongs to a corresponding fine classification category and the determining, according to the probability that the target image belongs to the corresponding fine classification category, of the classification result of the target image comprise: performing normalization processing on each of the third dot product processing results to obtain third normalized results; determining the third normalized results as the probability that the target image belongs to the corresponding fine classification category; and determining, as the classification result of the target image, a fine classification category corresponding to a maximum probability or a fine classification category corresponding to a target probability greater than a preset probability threshold.
14. A training device of an image classification model, comprising: an acquisition module configured to acquire n sample images and M texts from a training sample set; wherein the training sample set comprises N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of a target category, the texts are used for characteristic description of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; a training module configured to input the n sample images and the M texts into a to-be-trained model for training processing to obtain image classification results of each of the sample images; wherein the to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is obtained based on training of multiple image-text pairs; and the acquisition module is configured to acquire an image classification model from the to-be-trained model if a preset training end condition is met according to the image classification results; wherein the image classification model is used for classification processing of an input target image to be classified to obtain a target fine classification category to which the target image belongs under the target category; and wherein the to-be-trained model comprises a fusion module configured to perform correlation learning processing based on spliced features to obtain a correlation learning result, so as to determine the classification result of each of the sample images according to the correlation learning result, wherein the spliced features are features obtained by splicing target image features of the n sample images and target text features of the M texts. The determining of the classification result of each sample image according to the correlation learning result comprises: For each sample image, a target correlation learning result of the sample image about each fine classification category is obtained from the correlation learning result; Each target correlation learning result is dot product processed with a preset classification vector to obtain a second dot product processing result; According to the second dot product processing result, the probability of the sample image belonging to the corresponding fine classification category is determined, and the classification result of the sample image is determined according to the probabilities corresponding to the sample image.
15. An image classification device, comprising: an acquisition module that acquires a target image to be classified; a classification module that classifies the target image by using a pre-trained image classification model to obtain a classification result of the target image; the classification result represents a target fine classification category to which the target image belongs among a plurality of fine classification categories included in a target category; wherein the image classification model is obtained by training a to-be-trained model using n sample images and M texts acquired from a training sample set; the training sample set includes N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of the target category, and the texts are used to describe the characteristics of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; the to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is trained based on a plurality of image-text pairs; wherein the image classification model comprises a fusion module configured to perform correlation learning based on a spliced feature to obtain a correlation learning result, and determine the classification result of the target image according to the correlation learning result; the spliced feature is obtained by splicing a target image feature of the target image and a target text feature acquired; The determining of the classification result of the target image according to the correlation learning result comprises: obtaining a target correlation learning result of the target image about each fine classification category from the correlation learning result; dot product processing each target correlation learning result with a preset classification vector to obtain a third dot product processing result; According to the third dot product processing result, the probability of the target image belonging to the corresponding fine classification category is determined, and the classification result of the target image is determined according to the probability of the target image belonging to the corresponding fine classification category.
16. A training device of an image classification model, comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to: acquire n sample images and M texts from a training sample set; the training sample set includes N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of a target category, and the texts are used to describe the characteristics of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; input the n sample images and the M texts into a to-be-trained model for training processing to obtain an image classification result of each sample image; The to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is trained based on a plurality of image-text pairs; If it is determined according to the image classification result that a preset training end condition is met, an image classification model is obtained from the to-be-trained model; The image classification model is used for classifying an input target image to be classified, to obtain a target fine classification category to which the target image belongs under the target category; The to-be-trained model comprises a fusion module, the fusion module is used for performing correlation learning processing based on spliced features to obtain a correlation learning result, and the correlation learning result is used to determine a classification result of each sample image; The determination of the classification result of each sample image according to the correlation learning result comprises: For each sample image, a target correlation learning result of the sample image with respect to each fine classification category is obtained from the correlation learning result; Each target correlation learning result is dot multiplied with a preset classification vector to obtain a second dot multiplication result; According to the second dot multiplication result, a probability that the sample image belongs to a corresponding fine classification category is determined, and a classification result of the sample image is determined according to each probability corresponding to the sample image.
17. An image classification device, comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to: obtain a target image to be classified; classify the target image by using a pre-trained image classification model to obtain a classification result of the target image; the classification result represents a target fine classification category to which the target image belongs among a plurality of fine classification categories included in a target category; The image classification model is obtained by training a to-be-trained model using n sample images and M texts obtained from a training sample set; the training sample set comprises N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of the target category, and the texts are used to describe the characteristics of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; the to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is obtained by training a plurality of image-text pairs; The image classification model comprises a fusion module, the fusion module is used for performing correlation learning processing based on spliced features to obtain a correlation learning result, and the correlation learning result is used to determine the classification result of the target image, and the spliced features are features obtained by splicing target image features of the target image and target text features obtained; The determination of the classification result of the target image according to the correlation learning result comprises: A target correlation learning result of the target image with respect to each fine classification category is obtained from the correlation learning result; perform dot product processing on each of the target correlation learning result and a preset classification vector to obtain a third dot product processing result; According to the third dot product processing result, determine the probability that the target image belongs to the corresponding fine classification category, and determine the classification result of the target image according to the probability that the target image belongs to the corresponding fine classification category.
18. A computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions, when executed by a processor, implement the following processes: Obtain n sample images and M texts from a training sample set; the training sample set includes N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of a target category, and the texts are used to describe the characteristics of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; training processing in the trained model, to obtain an image classification result of each of the sample images; The to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is trained based on a plurality of image-text pairs; If it is determined that the preset training end condition is met according to the image classification result, an image classification model is obtained from the to-be-trained model; The image classification model is used for classifying an input target image to be classified to obtain a target fine classification category to which the target image belongs under the target category; The to-be-trained model includes a fusion module, which is used for correlation learning based on spliced features to obtain a correlation learning result, so as to determine the classification result of each sample image according to the correlation learning result; the spliced features are obtained by splicing target image features of the n sample images and target text features of the M texts; The determination of the classification result of each sample image according to the correlation learning result includes: For each sample image, obtain a target correlation learning result of the sample image with respect to each fine classification category from the correlation learning result; Perform dot product processing on each of the target correlation learning result and a preset classification vector to obtain a third dot product processing result; According to the third dot product processing result, determine the probability that the target image belongs to the corresponding fine classification category, and determine the classification result of the target image according to the probability that the target image belongs to the corresponding fine classification category.
19. A computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions, when executed by a processor, implement the following processes: Obtain a target image to be classified; Classify the target image by using a pre-trained image classification model to obtain a classification result of the target image; the classification result represents a target fine classification category to which the target image belongs among a plurality of fine classification categories included in a target category; The image classification model is obtained by training a to-be-trained model using n sample images and M texts obtained from a training sample set; the training sample set includes N sample images and the M texts, the M texts correspond one-to-one to M fine classification categories of a target category, and the texts are used to describe the characteristics of the corresponding fine classification categories; n, N and M are integers greater than 1, and n is less than N; The to-be-trained model is constructed based on a pre-trained model, and the pre-trained model is trained based on a plurality of image-text pairs; The image classification model comprises a fusion module, the fusion module is configured to perform correlation learning processing based on spliced features to obtain a correlation learning result, and the classification result of the target image is determined according to the correlation learning result, and the spliced features are features obtained by splicing the target image features of the target image and the obtained target text features; The classification result of the target image is determined according to the correlation learning result, comprising: Obtaining target correlation learning results of the target image with respect to each of the fine classification categories from the correlation learning result; Dot product processing is performed on each of the target correlation learning results and a preset classification vector to obtain a third dot product processing result; According to the third dot product processing result, the probability that the target image belongs to the corresponding fine classification category is determined, and the classification result of the target image is determined according to the probability that the target image belongs to the corresponding fine classification category.
Citation Information
Patent Citations
Image classification method and device, readable medium and electronic equipment
CN114511744A
Image classification method and apparatus, storage medium, and electronic device
WO2021138911A1