A multi-modal few-shot image classification method and device
By performing text semantic expansion and meta-learning strategy training on the data set, multimodal data sets are generated, which solves the problem of insufficient generalization ability of models in new categories in the existing technology, and achieves more efficient image classification performance.
Patent Information
- Application Number
- CN202411967091.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing few-sample learning methods have limitations in utilizing semantic information, resulting in insufficient generalization and distinction capabilities of the model in new categories, and it is impossible to effectively use limited data for learning and classification.
By augmenting text semantics of the category names in the dataset, a richer text description is generated, corresponding annotation vectors are generated for each image, and a multimodal data set is formed by combining image data, the classification model is pre-trained, and the model is further trained through meta-learning strategies to improve the model's generalization ability on a small number of samples.
It improves the performance of the model when understanding and generalizing new categories, enhances the distinction and generalization ability between categories, can more effectively use limited data for learning and classification, and improves the accuracy of small sample image classification.
Smart Images

Figure CN119963889B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software algorithms, and in particular, to a multi-modal few-shot image classification method and apparatus. Background Art
[0002] In recent years, with the development and maturity of Artificial Intelligence (AI) technology, its application in the field of image processing has become increasingly widespread. In particular, remarkable achievements have been made in the research of complex vision tasks such as Few-Shot Learning (FSL). Few-shot learning aims to quickly recognize new categories through limited sample data, which is an important and challenging task in the field of computer vision.
[0003] Traditional few-shot learning methods usually rely on prior knowledge obtained from large-scale base datasets. These methods mainly quickly adapt to new tasks through fixed embedding functions or parameter initialization. However, when dealing with rare samples, these methods face the following problems: current few-shot learning methods have limitations in utilizing semantic information. Some methods only rely on simple class names as text embeddings, while ignoring the semantic relationships between classes and richer prior knowledge. For example, class names usually only provide limited word-level information and fail to deeply utilize the semantic details and associations contained in the context or long texts. This deficiency limits the performance of the model in understanding and generalizing new categories, making it insufficient in the ability to distinguish between categories and generalization ability. This limitation makes the existing methods insufficient in the generalization ability and discrimination ability of new categories. Summary of the Invention
[0004] The present invention provides a multi-modal few-shot image classification method and apparatus to solve the defect of insufficient performance in the generalization ability and discrimination ability of new categories in the prior art, realize better generalization on a small number of samples, be able to more effectively utilize limited data for learning and classification, and improve the performance in understanding and generalizing new categories. The technical solutions proposed by the present invention are as follows:
[0005] In a first aspect, the present invention provides a multi-modal few-shot image classification method, including:
[0006] Obtain a pre-established dataset; wherein, the dataset includes multiple image data pairs, and each image data pair includes an image and a class name corresponding to the content of the image;
[0007] Perform text semantic augmentation on the class names in the dataset, and generate corresponding annotation vectors for each image based on the augmented text description to obtain a multi-modal dataset;
[0008] Pre-train a pre-established classification model based on the multi-modal dataset to obtain a pre-trained classification model;
[0009] Train the pre-trained classification model through a meta-learning strategy to obtain an image classification model;
[0010] Obtain a target image to be classified, and input the target image into the image classification model to obtain an image classification result.
[0011] Optionally, the annotation vector corresponding to the image is generated in the following manner:
[0012] Extract the definition information of the category name from a preset database, and use a pre-established large language model to perform text semantic expansion on the semantic information to obtain an expanded text description;
[0013] Use a tokenizer to decompose the expanded text description to obtain a sub-word token sequence;
[0014] Map each sub-word token in the sub-word token sequence to an identity identifier in a vocabulary to obtain a tokenization result;
[0015] Convert the tokenization result into a vector, and assign a sub-word weight to each sub-word in the vector to obtain the annotation vector.
[0016] Optionally, the sub-word weight is determined in the following manner:
[0017] Obtain the number of documents containing the sub-word in a corpus and the total number of documents in the corpus;
[0018] Determine the inverse document frequency according to the number of documents containing the sub-word and the total number of documents in the corpus;
[0019] Determine the word frequency of the sub-word in each document in the corpus;
[0020] Obtain the lengths of the documents in the corpus, and determine the average document length in the corpus according to the lengths of the documents in the corpus;
[0021] For each document in the corpus, determine the weight of the sub-word in the document according to the inverse document frequency, the word frequency of the sub-word in the document, the average document length, and the total number of documents in the corpus;
[0022] Determine the sub-word weight according to the weights of the sub-word in each document.
[0023] Optionally, the classification model includes a feature extractor and a classifier;
[0024] The pre-training of the pre-established classification model based on the multi-modal dataset to obtain a pre-trained classification model includes:
[0025] For each image in the multimodal dataset, divide the image into multiple sub-blocks to obtain the segmented image;
[0026] Input the segmented image into the feature extractor, and calculate the feature representation of each position in the segmented image through the multi-head self-attention mechanism;
[0027] Determine the feature representation corresponding to the image according to the feature representation of each position;
[0028] Determine the first loss according to the feature representations and class labels corresponding to the images in the multimodal dataset;
[0029] Optimize the feature extractor and the classifier by minimizing the first loss to obtain a pre-trained classification model.
[0030] Optionally, the feature extractor includes a text encoder, a first multi-layer perceptron, a second multi-layer perceptron, and a vision encoder; training the pre-trained classification model through a meta-learning strategy to obtain an image classification model includes:
[0031] Obtain a training set, randomly select a first number of categories from the training set, and randomly select a second number of samples with true labels for each category to obtain a support set for the task;
[0032] For each sample in the support set, use a pre-trained language model to extract the semantic features of the sample, and convert the semantic features to the same representation space as the true label through the first multi-layer perceptron to obtain the converted semantic features; wherein, the text semantic features are generated by combining learnable prompt vectors with category names;
[0033] In the training stage of meta-learning, use the vision encoder to extract the first visual features of the sample, and convert the first visual features to the same representation space as the true label through the second multi-layer perceptron to obtain the converted visual features;
[0034] Generate a global feature vector based on the converted semantic features and the converted visual features, and input the global feature vector into the classifier to obtain the original prediction scores for each category;
[0035] Calculate the second loss according to the original prediction scores and true labels of each category corresponding to each sample, and optimize the parameters of the feature extractor according to the second loss to obtain an image classification model.
[0036] Optionally, the method further includes:
[0037] Generate a query set for the task;
[0038] For each sample in the query set, use the vision encoder to extract the second visual features of the sample;
[0039] Obtain equi - length learnable vectors, fuse the second visual features of the samples and the equi - length learnable vectors to obtain the fused features;
[0040] Calculate the similarity between the fused features and the class prototypes of the support set to obtain the original prediction values for each class;
[0041] According to the original prediction values, select the class with the highest score as the prediction result;
[0042] For each sample in the query set, compare its prediction result with the true class and calculate the model evaluation metrics;
[0043] If the model evaluation metrics do not meet the preset requirements, retrain the pre - trained classification model.
[0044] In a second aspect, the present invention also provides a multi - modal few - shot image classification device, including the following modules:
[0045] A data acquisition module for acquiring a pre - established data set; wherein, the data set includes multiple image data pairs, and each image data pair includes an image and a class name corresponding to the content of the image;
[0046] A semantic expansion module for performing text semantic expansion on the class names in the data set, and generating corresponding annotation vectors for each image based on the expanded text description to obtain a multi - modal data set;
[0047] A first training module for pre - training a pre - established classification model based on the multi - modal data set to obtain a pre - trained classification model;
[0048] A second training module for training the pre - trained classification model through a meta - learning strategy to obtain an image classification model;
[0049] An image classification module for acquiring a target image to be classified, and inputting the target image into the image classification model to obtain an image classification result.
[0050] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, it implements the multi - modal few - shot image classification method as described in the first aspect above.
[0051] In a fourth aspect, the present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multi - modal few - shot image classification method as described in the first aspect above.
[0052] In a fifth aspect, the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the multi-modal few-shot image classification method as described in the first aspect above.
[0053] Based on the above technical solutions, the beneficial effects of the present invention compared with the prior art are as follows:
[0054] The multi-modal few-shot image classification method and device provided by the present invention generate richer text descriptions by performing text semantic augmentation on the class names in the dataset, enabling the model to capture the semantic relationships between classes and broader prior knowledge. Based on the augmented text descriptions, corresponding annotation vectors are generated for each image. These annotation vectors are the text representations of the image content, and together with the image data, they constitute a multi-modal dataset. This representation method enables the model to utilize text information to enhance the understanding of the image content while learning the image features. The classification model is pre-trained using the pre-processed multi-modal dataset. This process allows the model to learn the association between the image and the annotation vector, as well as the feature representation in the image. Since the model receives both image and text information as input, it can better utilize the complementarity between these two modalities to improve the classification performance. On the basis of pre-training, the model is further trained through a meta-learning strategy to improve the generalization ability of the model on a small number of samples, enabling it to more effectively utilize limited data for learning and classification, and enhancing the performance of the model in understanding and generalizing new classes. The present invention can make full use of long text information, not only focusing on the optimization of visual features, but also fully exploring the potential in semantic priors, improving the accuracy of few-shot image classification, and having great application prospects.
[0055] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, the claims, and the drawings.
[0056] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1It is a schematic flowchart of the multi-modal few-shot image classification method provided by the present invention.
[0059] Figure 2 It is a schematic diagram of the semantic augmentation process provided by the present invention.
[0060] Figure 3 It is a schematic diagram of the label generation method provided by the present invention.
[0061] Figure 4 It is a schematic structural diagram of the image classification model provided by the present invention.
[0062] Figure 5 It is a schematic structural diagram of the multi-modal few-shot image classification device provided by the present invention.
[0063] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0064] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0065] The following is combined with Figures 1 - 5 Describe the multi-modal few-shot image classification method and device of the present invention.
[0066] In order to make full use of the prior information in the semantics and further improve the accuracy of image classification, the present invention provides a multi-modal few-shot image classification method, which has high reliability and can improve the performance of the few-shot image classification algorithm. Referring to Figure 1 As shown, the multi-modal few-shot image classification method includes the following:
[0067] Step S110, obtain a pre-established data set; wherein, the data set includes multiple image data pairs, and each image data pair includes an image and a category name corresponding to the content of the image.
[0068] Collect and organize a dataset containing multiple pairs of image data. Specifically, collect image data from multiple sources (such as public datasets, web crawlers, self-built datasets, etc.). Ensure the diversity and representativeness of the image data, covering images in different categories, different scenarios, and different lighting conditions. Assign an accurate class name to each image. This can be done through manual annotation, automatic annotation (such as using a pre-trained classification model for prediction), or a combination of both. Organize the collected image data and class names into a dataset in a unified format for subsequent processing.
[0069] Specifically, each pair of image data consists of an image I and its corresponding class name T. For example, in the MiniImagnet dataset, in each pair of data samples (I, T), I is a scene image or a photo containing an object, and T is the corresponding class name. For example, the class name T corresponding to a dog is "A class of dog". The loaded pairs of image data can cover a wide range of scenarios and categories, thus providing a solid data foundation for the generalization ability of the model. After being loaded, the image I will be further processed to a fixed size (224×224) to meet the requirements of the model input. In addition, to enhance the diversity of the data and the robustness of the model, data augmentation operations will be performed on the image data, such as random cropping, rotation, color jitter, and horizontal flipping. These operations not only make the data more abundant but also effectively reduce the risk of overfitting. This dataset is the basis for all subsequent steps and is crucial for training a high-quality image classification model.
[0070] Step S120: Perform text semantic augmentation on the class names in the dataset, and generate corresponding annotation vectors for each image based on the augmented text description to obtain a multi-modal dataset.
[0071] Generally, in real life, it is usually very difficult to have a detailed text description of a certain category of objects. Considering the richness of semantics, the data preprocessing part of the present invention includes three parts: loading of image-text data, semantic augmentation, and label generation, aiming to provide high-quality input for subsequent model training through the generation of semantic cues and the optimization of annotation methods.
[0072] Load the image-text data pairs from the dataset. Each sample contains an image I and a class name T corresponding to the image content. Perform text semantic augmentation on the class names in the dataset to increase the richness of the text information and classification specificity, enabling the model to learn from finer-grained semantic information and improving the model's generalization ability and understanding ability. This step can use natural language processing (NLP) techniques, such as word embeddings, context embeddings, or large language models, to generate richer and more comprehensive text descriptions related to the class names. These augmented text descriptions not only contain the direct meaning of the class names but also cover other information related to them, such as attributes, context, or related concepts. For example, word embedding techniques (such as Word2Vec, GloVe, etc.) can be used to convert the class names into vector representations to capture their semantic information. Context embedding techniques (such as BERT, GPT, etc.) or querying online knowledge bases can also be utilized to obtain additional information related to the class names, such as attributes, context, related concepts, etc. Integrate this information into richer text descriptions and generate one or more relevant text descriptions for each image.
[0073] Based on the augmented text descriptions, generate corresponding annotation vectors for each image in the dataset. These annotation vectors are the text representations of the image content and will be used in subsequent classification tasks. The process of generating annotation vectors can adopt text vectorization techniques. The generation of annotation vectors enables image data to be compared and matched with text data in the same representation space.
[0074] Step S130: Pre-train a pre-established classification model based on the multimodal dataset to obtain a pre-trained classification model.
[0075] Use the multimodal dataset (i.e., the dataset containing images and annotation vectors) to pre-train the classification model. This classification model can be a deep learning model, such as a convolutional neural network (CNN) or a more complex Transformer architecture. Through pre-training, let the model learn the association between images and annotation vectors, as well as the feature representations in the images.
[0076] Specifically, select a deep learning model as the classification model, such as a convolutional neural network (CNN), Transformer architecture, etc. Perform necessary preprocessing operations on the model, such as adjusting the input dimensions, initializing the weights, setting hyperparameters, etc. Input the multimodal dataset into the model for training, allowing the model to learn the association between images and annotation vectors and the feature representations in the images. Strategies such as cross-validation and early stopping can be adopted during the training process to prevent overfitting.
[0077] Step S140: Train the pre-trained classification model through a meta-learning strategy to obtain an image classification model.
[0078] On the basis of pre-training, the classification model is further trained through a meta-learning strategy. The meta-learning strategy is an optimization strategy that aims to improve the generalization ability of the model on a small number of samples. In meta-learning training, the model is exposed to multiple different classification tasks, each of which contains a small number of labeled samples. Multiple subsets are randomly sampled from the multi-modal dataset as different classification tasks. Each task contains a small number of labeled samples (such as several images per category) and corresponding annotation vectors. Training is carried out on each task, and the hyperparameters or weights are adjusted according to the performance of the model. By continuously iterating this process, the model can learn how to more effectively utilize limited data for learning and classification.
[0079] Step S150: Obtain the target image to be classified, and input the target image into the image classification model to obtain an image classification result.
[0080] Apply the trained image classification model to new images to be classified. Obtain the images to be classified from the actual application scenario or user input. Input the images to be classified into the model, and the model will output a class name as the classification result. These results are obtained based on the degree of matching between the image content and the feature representations learned during training.
[0081] The multi-modal few-shot image classification method provided by the present invention first performs text semantic augmentation on the class names in the dataset to generate richer text descriptions, enabling the model to capture the semantic relationships between classes and broader prior knowledge. Based on the augmented text descriptions, corresponding annotation vectors are generated for each image. These annotation vectors are the text representations of the image content, and together with the image data, they constitute a multi-modal dataset. This representation method enables the model to utilize text information to enhance the understanding of image content while learning image features. Use the preprocessed multi-modal dataset to pre-train the classification model. This process allows the model to learn the association between images and annotation vectors, as well as the feature representations in the images. Since the model receives both image and text information as input, it can better utilize the complementarity between these two modalities to improve the classification performance. On the basis of pre-training, the model is further trained through a meta-learning strategy to improve the generalization ability of the model on a small number of samples, enabling it to more effectively utilize limited data for learning and classification.
[0082] Through text semantic augmentation and the generation of annotation vectors, the method of the present invention can capture deeper associations between images and class names. This helps the model to more accurately identify the content of images in classification tasks, thereby improving the accuracy of classification. The model can utilize richer class description information, which includes not only the direct meaning of the class name but also related attributes, contexts, or related concepts. This mechanism significantly enhances the performance of the model in understanding and generalizing new classes, strengthening its advantages in class discrimination ability and generalization ability. The use of the meta-learning strategy enables the model to better generalize on a small number of samples. This is of great significance for the small-sample learning problem in practical application scenarios because it allows the model to still exhibit good performance with limited training data. The method combines image data and text data to achieve multi-modal learning. This cross-modal information fusion helps to improve the robustness and adaptability of the model, enabling it to process more complex and diverse input data.
[0083] Traditional few-shot learning methods are prone to being interfered by background impurities when dealing with scarce samples, resulting in the deviation of feature extraction from the actual target and affecting the accurate expression of class features. The present invention indirectly enhances the model's ability to distinguish target features from background interference features by introducing multi-modal learning (the combination of images and text) and performing text semantic augmentation on class names in the preprocessing stage. By combining richer semantic information, the model can more accurately focus on target features when understanding image content, thereby reducing the impact of background interference to a certain extent.
[0084] In an optional embodiment, the annotation vector corresponding to the image in step S120 is generated in the following manner:
[0085] S1201. Extract the definition information of the class name from the preset database, and use the pre-established large language model to perform text semantic augmentation on the semantic information to obtain an augmented text description.
[0086] Obtain class names related to the image category (such as "cat", "dog", etc.) from the preset database. Extract the detailed definition information of these class names, which includes descriptions, attributes, common contexts, etc. Use a pre-trained large language model (such as BERT, GPT, etc.) to perform text semantic augmentation on the extracted semantic information to generate a more detailed and richer text description to capture the deep meaning and context information of the class name.
[0087] Such as Figure 2As shown, the first step in text augmentation is to extract the basic definition of the target object from a knowledge base (such as WordNet). For example, for the class name Name: "dog", its definition Definition can be obtained from the knowledge base: "A member of the genus Canis (Canis familiaris) that has been domesticated by humans since prehistoric times; occurs in many breeds".
[0088] This definition provides the scientific classification and basic characteristics of the target class, but it is usually too simple and lacks in-depth information for the classification task. To supplement the semantic expression of the definition and make up for the deficiencies in the knowledge base, a large language model (such as llama) is further used to expand the basic definition and generate a more detailed and highly relevant description for the classification task. This expanded content not only focuses on the key characteristics of the target class (such as color, hair type, behavior pattern, body proportion, etc.), but also combines scientificity and readability to ensure that the description is comprehensive, accurate and meets the actual application requirements. That is Figure 2 in: {Definition} provides a general definition of {Name}. Rewrite and expand this definition to include scientifically accurate details that highlight distinctive traits critical for classification, Focus on features that differentiate {Name} from other similar categories, such as unique color patterns, body proportions, behaviors, and textures. Keep the description concise and focused, using only one paragraph".
[0089] For example, an extended description of "dog" might be: "A dog (Canis familiaris), a domesticated subspecies of the gray wolf (Canis lupus), is notable for its extraordinary diversity in size, coat textures, colors, and behavioral adaptations, reflecting its evolutionary history and selective breeding for specific functions." Such extended descriptions are optimized to highlight the core features required for the classification task while remaining concise and avoiding redundant information. The ultimate goal is to provide strong support for the classification system through efficient and accurate semantic expression, and to improve the overall performance and accuracy of the task.
[0090] S1202. Use a tokenizer to decompose the augmented text description to obtain a sequence of sub-word tokens.
[0091] Use a tokenizer (such as Text Tokenizer) to decompose the augmented text description into a sequence of sub-word tokens (tokens). The tokenizer can identify words, phrases, or sub-word units in the text, preparing for subsequent vectorization processing.
[0092] S1203. Map each sub-word token in the sequence of sub-word tokens to an identity identifier in the vocabulary to obtain the tokenization result.
[0093] Map each sub-word token in the obtained sequence of sub-word tokens to a unique identity identifier (i.e., word embedding vector) in a predefined vocabulary. This step is a key step in converting text information into a numerical representation, facilitating subsequent calculations and processing.
[0094] S1204. Convert the tokenization result into a vector and assign a sub-word weight to each sub-word in the vector to obtain the labeled vector.
[0095] The mapped tokenized results can be converted into vector form, which can be achieved through word embeddings technology. A sub-word weight is assigned to each sub-word in the vector to reflect its importance in the overall semantics. This can be achieved through TF-IDF, word frequency statistics, or other weight assignment methods. The weighted sub-word vectors are combined to form the final annotation vector. This annotation vector contains rich semantic information about the image category and can be used for subsequent image classification, retrieval, or recognition tasks.
[0096] The above process of generating the label vector is illustrated by a specific example:
[0097] In the label generation step, the augmented text description T is converted into annotation information for model training. First, the augmented text description T is tokenized into sub-words using a tokenizer (Clip's Tokenlizer). The role of the tokenizer is to break the text into a sequence of finer-grained sub-word tokens. For example, for the text "A black dog playing in the park", the tokenizer splits it into a sequence of sub-word tokens: ["A", "black", "dog", "playing", "in", "the", "park"].
[0098] Each sub-word token is mapped to a unique ID in the vocabulary, such as [42, 103, 200, 345, 15, 3, 500]. These tokens can not only represent the semantics of the original text but also have high flexibility to adapt to the semantic learning needs of new categories. Then, the tokenized results are further processed into a K-hot vector (K-Hot Vector) to represent the sub-words that appear in the text. When generating the K-hot vector, a zero vector with a dimension of the vocabulary size V is initialized, and according to the position of the sub-word token, the value at the corresponding index is set to 1. For example, if the IDs corresponding to the sub-word sequence are [42, 103, 200], the generated K-hot vector, i.e., the above annotation vector, is [0, 0,..., 1,..., 0, 1,..., 1, 0,..., 0]. This vectorized annotation method can efficiently represent the semantic content in the text and is convenient for the model to process.
[0099] The present invention expands text semantics through a large language model, which can capture the deep meaning and context information of category names, thereby improving the model's ability to understand image content. Using a tokenizer and a vocabulary to convert text information into a numerical representation can ensure the accuracy and consistency of the annotation vectors. By assigning weights to subwords, key information can be further highlighted, improving the accuracy of annotation. The rich semantic information and accurate annotation vectors can help the model better generalize to unseen images. This is crucial for improving the performance of image classification, retrieval, or recognition tasks. The generated annotation vectors can be combined with image feature vectors to support multimodal learning, thereby further improving the performance and accuracy of the model.
[0100] In an optional embodiment, in natural language, the contributions of different subwords to semantics are not equal. Therefore, corresponding subword weights need to be assigned to each subword. The subword weights in the above step S1204 are determined by the following method:
[0101] S12041. Obtain the number of documents containing the subword in the corpus and the total number of documents in the corpus.
[0102] Obtain corpus statistics, that is, the number of documents containing the subword and the total number of documents in the corpus. First, traverse the entire corpus and count how many documents each subword appears in. This reflects the universality or rarity of the subword. At the same time, record how many documents there are in total in the corpus. This is the basis for calculating the inverse document frequency (IDF).
[0103] S12042. Determine the inverse document frequency according to the number of documents containing the subword and the total number of documents in the corpus.
[0104] Use the formula IDF(t) = log(N / DF(t) + 1), where IDF(t) represents the inverse document frequency of subword t, N is the total number of documents in the corpus, and DF(t) is the number of documents containing subword t. Adding 1 is to avoid the denominator being 0. IDF measures the rarity of the subword in the entire corpus.
[0105] S12043. Determine the term frequency of the subword in each document in the corpus.
[0106] For each document in the corpus, count the number of times each subword appears in the document, that is, the term frequency (TF). This reflects the importance of the subword in a specific document.
[0107] S12044. Obtain the lengths of the documents in the corpus, and determine the average document length in the corpus according to the lengths of the documents in the corpus.
[0108] Traverse the corpus, calculate the length of each document (i.e., the total number of subwords in the document), and then find the average of these lengths. The average document length is used in subsequent weight calculations to account for differences in document length.
[0109] S12045. For each document in the corpus, determine the weight of a subword in the document according to the inverse document frequency, the word frequency of the subword in the document, the average document length, and the total number of documents in the corpus.
[0110]
[0111] Among them, represents the weight of the subword in the document , TF(t, d) represents the word frequency of the subword t in the document d, ∣d∣ represents the length of the document d, and avgdl represents the average document length in the corpus. and are adjustment factors, which are set according to specific situations. For example, can be set to a value between 1.2 and 2, can be set to 0.75. This formula takes into account both the universality and rarity of subwords through , and also takes into account the importance of subwords in specific documents through , while also considering differences in document length.
[0112] S12046. Determine the subword weights according to the weights of subwords in each document.
[0113] For each subword, the average or weighted average of its weights in all documents can be calculated as the final weight of the subword.
[0114] By introducing weights, long-tail categories or low-frequency subwords receive more attention during training, thereby improving the model's ability to recognize few-shot categories. As Figure 3 shown, for the text "A black dog playing in the park", the weights of the subwords "dog", "black", and "park" are calculated to be 2.6, 1.8, and 1.5 respectively. The final generated annotation vector will be adjusted according to these weights to make it closer to the true semantic distribution. The annotation vector y = [0,…, 2.6(dog),…, 1.8(black),…, 1.5(park),.., 0]. At this time, the weight of the important subword "dog" dominates, while the weights of frequently occurring but less informative subwords (such as "the", "in") are close to 0.
[0115] By combining the inverse document frequency, the sub - word frequency in the document, and the document length information, the present invention can more accurately reflect the importance of sub - words in the corpus and specific documents. This helps to more accurately utilize sub - word information in subsequent text processing or machine learning tasks. Considering the difference in document length can reduce the weight bias caused by different document lengths, thereby improving the stability and robustness of the model. The weight determination method is based on corpus statistical information, so it is easy to integrate with other text processing or machine learning algorithms. At the same time, it can also be extended as needed to consider more text features or context information.
[0116] In an optional embodiment, the classification model in the above - mentioned step S130 includes a feature extractor and a classifier, and the classifier can be a linear classification head. The goal of model pre - training is to classify a large - scale multimodal dataset , and optimize the feature extractor f and the linear classification head [W, b] by minimizing the loss.
[0117] The pre - training of the pre - established classification model based on the multimodal dataset in the above - mentioned step S130 to obtain a pre - trained classification model includes:
[0118] 1301. For each image in the multimodal dataset, divide the image into multiple sub - blocks to obtain the segmented image.
[0119] For each image in the multimodal dataset, to enhance the model's ability to model visual features, image segmentation processing is first performed. The purpose of this step is to divide the image into multiple smaller sub - blocks so that the subsequent feature extraction process can capture the information in the image more meticulously. The segmented image sub - blocks should be able to fully reflect the content and structure of the original image while maintaining a certain degree of independence to facilitate the independent processing of each sub - block by the feature extractor.
[0120] 1302. Input the segmented image into the feature extractor Vision Transformer, calculate the feature representation of each position in the segmented image through the multi - head self - attention mechanism, and determine the feature representation corresponding to the image according to the feature representation of each position.
[0121] Input the segmented image into the feature extractor. During the feature extraction process, the multi - head self - attention mechanism is used to calculate the feature representation of each position in the segmented image. The multi - head self - attention mechanism can capture the correlation information between different positions in the image, thereby generating a richer feature representation. Through the processing of the feature extractor, each image sub - block will obtain a corresponding feature vector, and these feature vectors together constitute the feature representation of the image.
[0122] Based on the feature representations at each position, a certain aggregation strategy (such as average pooling, max pooling, etc.) is used to determine the feature representation corresponding to the entire image. The choice of the aggregation strategy should be able to fully reflect the information of each sub-block in the image while maintaining the simplicity and effectiveness of the feature representation. Taking average pooling as an example, the feature representation of the image is determined by the following formula:
[0123]
[0124] where, the feature representation generated by the feature extractor, is the feature at position i in the L-th layer of the Transformer. is the number of sub-blocks in the segmented image.
[0125] 1303. Determine the first loss according to the feature representations and class labels corresponding to each image in the multi-modal dataset.
[0126] Calculate the first loss according to the feature representations and class labels corresponding to each image in the multi-modal dataset. This step usually involves a classifier that predicts the class label of the image based on the feature representation. The first loss can be calculated using common loss functions such as cross-entropy loss, mean squared error loss, etc. The choice of the loss function should be able to accurately reflect the difference between the prediction result and the true label, thereby guiding the optimization process of the model.
[0127]
[0128] where, the first loss, represents the input image, represents the multi-modal dataset, represents the true class corresponding to the image, that is, the above-mentioned annotation vector. the weights and biases of the classifier corresponding to the class of. the feature representation generated by the feature extractor. represents the true class score of, true class corresponding weights and biases, represents the exponential function. represents the true class exponential value of the score of. represents the sum of the exponential values of all class scores.
[0129] 1304. Optimize the feature extractor and the classifier by minimizing the first loss to obtain a pre-trained classification model.
[0130] Optimize the feature extractor and classifier by minimizing the first loss. This step can adopt optimization algorithms such as gradient descent to reduce the loss value by iteratively updating the parameters of the model. During the optimization process, hyperparameters such as the learning rate and batch size can be adjusted as needed to accelerate the convergence of the model and improve the performance of the model.
[0131] In the present invention, by dividing the image into multiple sub-blocks and processing them independently, the feature extractor can capture the information in the image more meticulously, thereby improving the accuracy and robustness of feature extraction. The multi-head self-attention mechanism can capture the correlation information between different positions in the image, enabling the model to better understand the overall structure and content of the image, thereby enhancing the generalization ability of the model. By minimizing the first loss to optimize the feature extractor and classifier, the model can achieve better performance on the training set. At the same time, due to the adoption of the pre-training strategy, the model can also adapt and converge faster in subsequent tasks. This process is not only applicable to image processing tasks but can also be extended to data processing tasks of other modalities. By constructing the corresponding feature extractor and classifier, the joint processing and analysis of multi-modal data can be realized.
[0132] In an optional embodiment, the present invention trains the model in an N-way K-shot task through a meta-learning strategy, enabling the model to have generalization ability. Refer to Figure 4 As shown, the feature extractor of the image classification model includes a text encoder, a first multi-layer perceptron, a second multi-layer perceptron, and a visual encoder; training the pre-trained classification model through the meta-learning strategy described in step S140 above to obtain an image classification model includes:
[0133] S1401. Obtain a training set, randomly select a first number of categories from the training set, and randomly select a second number of samples with true labels for each category to obtain the support set of the task.
[0134] N-way K-shot refers to a task setting where there are N different categories (way) and K samples (shot) in each category. This setting aims to simulate the scenario of scarce labeled data in the real world, thereby evaluating the generalization ability of the model under a small amount of labeled data. In the task, the model needs to accurately classify or identify new categories when only seeing a small number of labeled samples.
[0135] In the N-way K-shot setting, the task is divided into a Support Set and a Query Set. The Support Set contains N classes, with K samples in each class. The Query Set contains samples of the same classes as in the Support Set, but the samples themselves do not overlap, and is used to evaluate the performance of the model. Randomly select N classes from the training set, and then randomly select K samples with true labels from each class to form the Support Set S = {(xs, ys)}. xs represents the image, and ys represents the class name of the sample. At the same time, construct the Query Set Q = {(xq, yq)}, where xq is the sample to be classified, and yq is the true label that is invisible during testing and needs to be predicted by the model. The Support Set is used to train the model to learn the feature representations of new classes, while the Query Set is used to evaluate the classification performance of the model.
[0136] S1402. For each sample in the Support Set, use a pre-trained language model to extract the semantic features of the sample, and transform the semantic features to the same representation space as the true label through a first multi-layer perceptron to obtain the transformed semantic features; wherein, the text semantic features are generated by combining a learnable prompt vector and the class name.
[0137] When constructing the model, a same learnable prompt vector is preset. Specifically, use the phrase embedding in the language model (such as "a photo of a") as the initial value. This vector is in the form of a continuous and optimizable parameter, and is combined with the detailed class description to form a complete Prompt and input into the text encoder. Initialize the pre-trained classification model, including a Vision Encoder, a Text Encoder, and a Classifier. Set the parameters of the meta-learning task, such as N and K in N-way K-shot, where N and K represent the number of classes and the number of samples in each class in each task, respectively.
[0138] Before the model training starts, generate an initial learnable prompt vector for the samples in the Support Set. To enhance the expressive ability of the model, use the pre-trained language model, i.e., the above-mentioned text encoder g(⋅), to extract semantic features. By combining the learnable prompt vector and the class name, input into the pre-trained language model to generate the semantic feature g(y text )). This process enhances the expressive ability of the model and enables the model to understand the semantic information of the class name. The semantic feature g(y text ) is expressed as:
[0139] g(y text) = g(v1, v2, …, vL, [classname])
[0140] Among them, vi is the i-th element of the learnable prompt vector, classname is the class name, and L is the number of elements of the learnable prompt vector.
[0141] The above text encoder can adopt the Long-CLIP model, which is developed on the basis of the Contrastive Language–Image Pre-training (CLIP) model and supports long text input.
[0142] Use the visual encoder to extract the visual features g(x image ). The visual features reflect the visual content of the sample and are an important information source for classification tasks.
[0143] S1403. In the training stage of meta-learning, use the visual encoder to extract the first visual features of the sample, and convert the first visual features to the same representation space as the true label through the second multi-layer perceptron to obtain the converted visual features;
[0144] Convert the semantic features to the same representation space as the true label through the first multi-layer perceptron (MLP1) to obtain the converted semantic feature Z. Similarly, convert the visual features to the same representation space through the second multi-layer perceptron (MLP2) to obtain the converted visual feature W.
[0145] Z = MLP1(g(y text ))
[0146] W = MLP2(g(x image ))
[0147] S1404. Generate a global feature vector based on the converted semantic features and the converted visual features, and input the global feature vector into the classifier to obtain the original prediction scores logits for each class. Logits are unnormalized scores and represent the original measure of the confidence or probability for each class.
[0148] When fusing the converted semantic feature Z and visual feature W, simple multi-modal feature fusion can be adopted, such as the addition method.
[0149] S1405. Calculate the second loss based on the original prediction scores and true labels of each category corresponding to each sample, and optimize the parameters of the feature extractor according to the second loss to obtain an image classification model. Use the Softmax multi-label classification loss function to compare the logits output by the model with the K-hot labels and calculate the second loss, such as the cross-entropy loss. The K-hot label indicates the K categories that a sample may belong to among N categories and is a sparse vector. Through the backpropagation algorithm, optimize the parameters of the feature extractor (including the text encoder, the first multi-layer perceptron, the second multi-layer perceptron, and the visual encoder) according to the calculated loss value to gradually improve the performance of the model in the N-way K-shot task.
[0150]
[0151] Among them, represents the cross-entropy loss function, which is used to measure the difference between the probability distribution predicted by the model and the true label distribution. The smaller the cross-entropy loss, the closer the probability distribution predicted by the model is to the true label distribution, that is, the better the performance of the model.
[0152] represents the original prediction score (the score without Softmax normalization) of the model for the th category. These outputs come from the last layer of the model, and there is a corresponding score for each category. In multi-label classification, the model outputs a score for each category, and these scores are then converted into probabilities through the Softmax function, so that the scores of each category are between 0 and 1, and the sum of the probabilities of all categories is 1.
[0153] The K-hot label represents the true label of the th category. In multi-label classification, the K-hot label means that there are K positions in the label vector that are 1 (indicating that the categories corresponding to these positions exist), and the remaining positions are 0 (indicating that the categories corresponding to these positions do not exist). Different from the one-hot encoding in single-label classification (only one position is 1), multi-label classification allows multiple positions to be 1.
[0154] represents the total number of categories, and the model needs to predict the existence or non-existence of categories.
[0155] represents the original output score (logits) of the model for the th category. is an index variable that traverses all possible categories = 1, 2, ..., 。
[0156] represents the Softmax function, which converts logits into a probability distribution. For each class , the Softmax function calculates the ratio to the sum of the exponents of all classes, thereby obtaining the probability of that class.
[0157] Through the backpropagation algorithm, the loss value is passed back to each layer of the model, and the gradient of each parameter is calculated. These gradients are used to update the parameters of the model, including the values of the learnable prompt vectors. In each iteration, according to the calculated gradients, optimization algorithms (such as SGD, Adam, etc.) are used to update the parameters of the model. These updates include the values of the learnable prompt vectors, making them gradually approach the optimal solution that can better represent the class features. As the training progresses, the loss value of the model will gradually decrease, and the values of the learnable vectors will also gradually stabilize. When the loss value reaches a sufficiently small threshold or no longer decreases significantly, it can be considered that the model has converged, and the learnable vectors obtained at this time are the final learned equal-length learnable vectors.
[0158] The present invention trains the model in the N-way K-shot task through a meta-learning strategy, enabling the model to learn how to quickly adapt to new classes and improving the generalization ability of the model. By combining text semantic features and visual features for multi-modal feature fusion, various information sources of the samples are fully utilized, improving the classification accuracy. By introducing a learnable prompt vector to combine with the class name to generate semantic features, the expression ability of the model is enhanced, enabling the model to better understand the semantic information of the class name. Using the Softmax multi-label classification loss function and the backpropagation algorithm to optimize the model parameters enables the model to gradually approach the optimal solution during training, improving the classification performance of the model.
[0159] In an optional embodiment, in order to evaluate the classification ability of the model on unseen classes, a query set needs to be generated. The method further includes:
[0160] S201. Generate a query set for the task.
[0161] The query set Q = {(xq, yq)} contains the same classes as the support set, but different samples, and each class can have multiple samples. These samples also include the image xq and the corresponding true label yq that is invisible during testing, which the model needs to predict.
[0162] S202. For each sample in the query set, use the visual encoder to extract the second visual feature of the sample.
[0163] Extract the second visual feature of each sample in the query set using the same visual encoder as during training (i.e., step S1403 above).
[0164] S203. Obtain an equal-length learnable vector, and fuse the second visual feature of the sample and the equal-length learnable vector to obtain a fused feature.
[0165] Obtain an equal-length learnable vector. The length of this vector is the same as the dimension of the semantic feature or visual feature after being transformed by a multi-layer perceptron, and it is used to fuse with the visual feature during the inference process. Fuse the second visual feature of each sample in the query set and the equal-length learnable vector to obtain a fused feature. This step is similar to the processing method of the samples in the support set during training.
[0166] S204. Calculate the similarity between the fused feature and the class prototype of the support set to obtain the original prediction value for each class; according to the original prediction value, select the class with the highest score as the prediction result. The vector of the support set obtains the corresponding class prototype through the following formula:
[0167]
[0168] where and are the transformed semantic feature and visual feature, is the finally output classification vector, that is, the above-mentioned class prototype, which is used to represent the probability that the sample belongs to each class, is the number of samples.
[0169] S207. For each sample in the query set, compare its prediction result with the true class, and calculate the model evaluation index; if the model evaluation index does not meet the preset requirements, retrain the pre-trained classification model.
[0170] For each sample in the query set, compare its prediction result with the true class, and calculate model evaluation indexes such as accuracy, recall rate, and F1 score. If the model evaluation index does not meet the preset requirements (such as the accuracy being lower than a certain threshold), it indicates that the performance of the model on the current task is not good, and the pre-trained classification model needs to be retrained through the above steps S1401 - S1405. When retraining, the hyperparameters of the model can be adjusted, the training data can be increased, the feature extraction method can be improved, etc., to improve the performance of the model.
[0171] The multi-modal few-shot image classification method based on long text provided by the present invention extracts the semantic features of category names by introducing a pre-trained language model, and performs multi-modal fusion in combination with visual features, thereby making full use of the semantic content in the long text information and improving the classification accuracy. The model is trained in the N-way K-shot task through a meta-learning strategy, enabling the model to learn how to quickly adapt to new categories and achieve good classification results even with only a small number of samples, enhancing the generalization ability of the model. During the inference process, the model can dynamically generate fusion features based on the input image and category name, and output the prediction values for each category. This flexible inference process enables the model to be applicable to different application scenarios and task requirements. By calculating the model evaluation metrics and determining whether retraining is needed, the performance of the model can be continuously optimized to make it more stable and reliable in practical applications.
[0172] The multi-modal few-shot image classification device provided by the present invention will be described below. The multi-modal few-shot image classification device described below can be mutually referred to the multi-modal few-shot image classification method described above.
[0173] The multi-modal few-shot image classification device provided by the present invention refers to Figure 5 as shown, and includes:
[0174] A data acquisition module 310, configured to acquire a pre-established data set; wherein, the data set includes a plurality of image data pairs, and each image data pair includes an image and a category name corresponding to the content of the image;
[0175] A semantic expansion module 320, configured to perform text semantic expansion on the category names in the data set, and generate a corresponding annotation vector for each image based on the expanded text description to obtain a multi-modal data set;
[0176] A first training module 330, configured to pre-train a pre-established classification model based on the multi-modal data set to obtain a pre-trained classification model;
[0177] A second training module 340, configured to train the pre-trained classification model through a meta-learning strategy to obtain an image classification model;
[0178] An image classification module 350, configured to acquire a target image to be classified, and input the target image into the image classification model to obtain an image classification result.
[0179] Figure 6 An example of a schematic physical structure diagram of an electronic device is shown in Figure 6As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the multi-modal few-shot image classification method.
[0180] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0181] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-modal few-shot image classification method provided by the above-mentioned various methods.
[0182] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the multi-modal few-shot image classification method provided by the above-mentioned various methods.
[0183] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0184] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-modal few-shot image classification method, characterized in that, Including: Obtain a pre-established dataset; wherein, the dataset includes multiple image data pairs, and each image data pair includes an image and a category name corresponding to the content of the image; Perform text semantic augmentation on the category names in the dataset, and generate corresponding annotation vectors for each image based on the augmented text description to obtain a multi-modal dataset; the annotation vector corresponding to the image is generated by the following method: extract the definition information of the category name from a preset database, and use a pre-established large language model to perform text semantic augmentation on the definition information to obtain the augmented text description; use a tokenizer to decompose the augmented text description to obtain a sub-word token sequence; map each sub-word token in the sub-word token sequence to an identity identifier in the vocabulary to obtain a tokenization result; convert the tokenization result into a vector, and assign a sub-word weight to each sub-word in the vector to obtain the annotation vector; the sub-word weight is determined by the following method: obtain the number of documents containing the sub-word in the corpus and the total number of documents in the corpus; determine the inverse document frequency according to the number of documents containing the sub-word and the total number of documents in the corpus; determine the word frequency of the sub-word in each document in the corpus; obtain the lengths of the documents in the corpus, and determine the average document length in the corpus according to the lengths of the documents in the corpus; for each document in the corpus, determine the weight of the sub-word in the document according to the inverse document frequency, the word frequency of the sub-word in the document, the average document length and the total number of documents in the corpus; determine the sub-word weight according to the weights of the sub-word in each document; Pre-train a pre-established classification model based on the multi-modal dataset to obtain a pre-trained classification model; Train the pre-trained classification model through a meta-learning strategy to obtain an image classification model; Obtain a target image to be classified, and input the target image into the image classification model to obtain an image classification result.
2. The multi-modal few-shot image classification method according to claim 1, wherein The classification model includes a feature extractor and a classifier; The pre-training of the pre-established classification model based on the multi-modal dataset to obtain a pre-trained classification model includes: For each image in the multi-modal dataset, divide the image into multiple sub-blocks to obtain the segmented image; Input the segmented image into the feature extractor, and calculate the feature representation of each position in the segmented image through a multi-head self-attention mechanism; Determine the feature representation corresponding to the image according to the feature representation of each position; Determine a first loss according to the feature representations and category labels corresponding to the images in the multi-modal dataset; Optimize the feature extractor and the classifier by minimizing the first loss to obtain a pre-trained classification model.
3. The multimodal few-shot image classification method according to claim 2, wherein The feature extractor includes a text encoder, a first multi-layer perceptron, a second multi-layer perceptron and a visual encoder; The training of the pre-trained classification model through a meta-learning strategy to obtain an image classification model includes: Obtain a training set, randomly select a first number of categories from the training set, and randomly select a second number of samples with true labels for each category to obtain a support set for the task; For each sample in the support set, use a pre-trained language model to extract the semantic features of the sample, and transform the semantic features into the same representation space as the true label through a first multi-layer perceptron to obtain the transformed semantic features; wherein, the text semantic features are generated by combining learnable prompt vectors with class names. In the training stage of meta-learning, use the visual encoder to extract the first visual features of the sample, and transform the first visual features into the same representation space as the true label through a second multi-layer perceptron to obtain the transformed visual features. Generate a global feature vector based on the transformed semantic features and the transformed visual features, and input the global feature vector into a classifier to obtain the original prediction scores for each category. Calculate a second loss based on the original prediction scores and true labels for each category corresponding to each sample, and optimize the parameters of the feature extractor according to the second loss to obtain an image classification model.
4. The multi-modal few-shot image classification method according to claim 3, wherein The method further includes: Generating a query set for the task. For each sample in the query set, use the visual encoder to extract the second visual features of the sample. Obtain an equal-length learnable vector, and fuse the second visual features of the sample with the equal-length learnable vector to obtain the fused features. Calculate the similarity between the fused features and the class prototypes of the support set to obtain the original prediction values for each category. According to the original prediction values, select the category with the highest score as the prediction result. For each sample in the query set, compare its prediction result with the true category and calculate the model evaluation metrics. If the model evaluation metrics do not meet the preset requirements, retrain the pre-trained classification model.
5. A multi-modal few-shot image classification device, characterized in that, Including: A data acquisition module for acquiring a pre-established data set; wherein, the data set includes a plurality of image data pairs, and each image data pair includes an image and a class name corresponding to the content of the image. A semantic augmentation module, which is used to perform text semantic augmentation on the category names in the dataset, and generate corresponding annotation vectors for each image based on the augmented text description to obtain a multi-modal dataset; the annotation vector corresponding to the image is generated in the following manner: extract the definition information of the category name from a preset database, and use a pre-established large language model to perform text semantic augmentation on the definition information to obtain the augmented text description; use a tokenizer to decompose the augmented text description to obtain a sub-word token sequence; map each sub-word token in the sub-word token sequence to an identity identifier in a vocabulary to obtain a tokenization result; convert the tokenization result into a vector, and assign a sub-word weight to each sub-word in the vector to obtain the annotation vector; the sub-word weight is determined in the following manner: obtain the number of documents containing the sub-word in the corpus and the total number of documents in the corpus; determine the inverse document frequency according to the number of documents containing the sub-word and the total number of documents in the corpus; determine the word frequency of the sub-word in each document in the corpus; obtain the lengths of the documents in the corpus, and determine the average document length in the corpus according to the lengths of the documents in the corpus; for each document in the corpus, determine the weight of the sub-word in the document according to the inverse document frequency, the word frequency of the sub-word in the document, the average document length and the total number of documents in the corpus; determine the sub-word weight according to the weights of the sub-word in each document. A first training module, which is used to pre-train a pre-established classification model based on the multi-modal dataset to obtain a pre-trained classification model. A second training module, which is used to train the pre-trained classification model through a meta-learning strategy to obtain an image classification model. An image classification module, which is used to obtain a target image to be classified, and input the target image into the image classification model to obtain an image classification result.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-modal few-shot image classification method according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-modal few-shot image classification method according to any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-modal few-shot image classification method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Multimodal data heterogeneous transformer-based asset recognition method, system, and device
US12236699B1