A Few-Shot Learning Method for Task-Adaptive Correlation Learning

By designing task-specific prompt words and comparative learning to enhance training features, the problems of insufficient semantic information and inaccurate features in the existing small sample learning methods are solved, and more efficient small sample learning and generalization capabilities are achieved.

CN119091248BActive Publication Date: 2025-06-17XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411136625.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-06-17
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

The existing small sample learning method lacks semantic information when dealing with category name encoding, resulting in misjudgment; at the same time, the training features are not processed, resulting in inaccurate results when querying in similar samples.

Method used

By designing task-specific prompt words, using large language models to optimize prompt words to improve the recognition performance of the CLIP model; at the same time, comparative learning is used to enhance the discriminantity of training features, and by designing linear layer network adapter and cross-entropy loss function, feature representation is optimized.

Benefits of technology

It significantly improves the learning ability and generalization ability of the model in small samples or even zero samples, reduces the dependence on large-scale annotation data sets, and enhances the adaptability and accuracy of the model in the face of new environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091248B_ABST
    Figure CN119091248B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image recognition and relates to a few-shot learning method for task-adaptive correlation learning, including: S1, using task-specific prompt words to assist the original prompts; S2, using contrastive learning to enhance the discriminability of training features; S3, using diverse image enhancement means to enhance the contrast of images; The present invention strengthens the generalization ability of the model through rich descriptions in natural language, enabling CLIP to handle even previously unseen categories or complex concepts with ease. Therefore, the method of the present invention not only reduces the dependence on large-scale labeled datasets, but also improves the learning ability of the model in few-shot or even zero-shot scenarios, enhances the adaptability and generalization ability of the model to new environments, and greatly simplifies the model deployment process, making it a powerful tool for solving a series of challenges in image recognition and description tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a few-shot learning method for task-adaptive correlation learning. Background Art

[0002] With the rapid development of deep learning in the field of images, the computer's recognition of images has approached or even surpassed human performance. Although deep learning has achieved remarkable results in computer vision fields such as object recognition, scene understanding, and semantic segmentation, there are also some key bottleneck problems. For example, existing deep learning models rely heavily on centralized large-scale labeled training data. When the labeled training data is scarce, the generalization ability of deep learning models will be significantly reduced, resulting in limited application scenarios. Different from the current deep learning models, humans can generally master new target concepts quickly with only a small amount of data. Inspired by this, researchers have proposed the concept of "few-shot learning", aiming to explore how to use the knowledge and experience summarized from existing samples to solve new problems when the labeled samples of new target categories are extremely few.

[0003] The few-shot learning method is to improve the test accuracy under relatively few training samples. The specific method is to extract the image features in the training dataset through the CLIP model (Contrastive Language-Image Pre-Training, a large multi-modal pre-training model proposed by openAI, which includes text and image encoders. Define the positive samples as image-text pairs of the same class, and the negative samples as image-text pairs other than that. Learn the corresponding relationship between the two modalities of image and text through contrastive learning. It can output the similarity between a given text and an image), store them in a cache model, and define these image features as keys. At the same time, convert the labels corresponding to the training images into one-hot labels, called values. For a test image (which can be called a query), predict its label in two ways. The first is to calculate its similarity with the training images. If it is similar to a certain image, it is considered that the test sample belongs to the class to which this image belongs. The specific calculation process is as Figure 1 shown in the horizontal process in the figure. By multiplying the test sample image features and the training sample image feature set, and then multiplying the label set, the corresponding prediction vector is obtained. The second is through the zero-shot inference of the CLIP model. First, encode the text features of all class names on this dataset to form W_c, and then multiply it with the test image features to get the similarity, that is, the prediction vector. Finally, add these two prediction vectors weighted and select the class index corresponding to the maximum logits as the predicted class, as Figure 1 shown.

[0004] However, the above few-shot learning method has the following defects:

[0005] Disadvantage 1: When this method uses CLIP to infer the categories of test samples, it uses the text embeddings of all category names in the dataset. However, only encoding the semantic information contained in the category names is too little, because some category names may be very similar or have multiple meanings, and they cannot accurately describe the samples of this category, which will cause misjudgment. For example, for a specific task, such as cat breed classification, all the data in the dataset are various cats, so the category names may be very similar, which will lead to very similar generated text features, making it difficult to distinguish different categories. Or for the classification of some plant leaves, the category names are all "such and such leaves", so the repeated appearance of the word "leaves" will also increase category confusion and thus increase the classification difficulty.

[0006] Disadvantage 2: This method simply stores the features of training samples in the cache model without processing these features. The training samples of few-shot tasks are particularly few, so the representation of each sample is crucial for the correct classification of test samples. When encountering visually very similar samples, the extracted features will undoubtedly be very similar. When the test sample queries the most similar sample in such a feature set, the obtained result is bound to be unable to guarantee accuracy. Summary of the Invention

[0007] To solve the above technical problems, the present invention provides the following technical solutions: A few-shot learning method for task-adaptive association learning, the method comprising the following steps:

[0008] S1. Use task-specific prompt words to assist the original prompt; specifically including:

[0009] S1-1. Design a template for this task, the template includes: Template 1, Template 2. Template 1 includes: Using the knowledge of experts in this field to generate more accurate prompt words for the category. Template 2 includes: Using vivid and intuitive descriptive words to directly describe the visual features presented by this category. By designing these two templates, it can effectively avoid the generated prompt words being general and containing abstract adjectives, so that CLIP cannot accurately understand and locate to the appropriate position in the joint embedding space (CLIP learns the image and text feature embeddings through contrastive learning, which is an aligned joint embedding space, that is, the text and images of the same category are aligned in the embedding space), and then find the most similar image features.

[0010] S1-2, input the two templates designed in step S1-1 into the large language model in sequence, that is, first write the category name into template one and input it into LLM, you will get one or more prompt words, and then further input template two into LLM in a dialogue way, you will get the optimized prompt words; traverse all the categories in the data set, you can get a series of task-specific prompt words, save them, and use them in subsequent model reasoning;

[0011] S2. Use contrastive learning to enhance the discriminability of training features; specifically including:

[0012] S2-1. Design a linear layer network adapter, the size of which is consistent with the transpose of the size of the feature set, and initialize the network with the feature set extracted by the image feature extractor of CLIP; input the features of the input image into the adapter to obtain an affinity matrix affinity, which represents the similarity matrix between the input features and the training features; normalize affinity to obtain affinity_norm, take this variable as the object of optimization, and design a contrastive learning loss, which aims to shorten the distance between positive samples and push the distance between negative samples away, so as to learn useful feature representations;

[0013] S2-2, calculating the similarity matrix between the eigenvectors by multiplying the eigenvector matrix with its transpose and dividing by a temperature parameter; the temperature parameter can control the scale of the obtained similarity, a lower temperature will make the differences in the similarity matrix more significant, while a higher temperature will make these differences more gradual;

[0014] S2-3. Create a Boolean mask matrix mask with True on the diagonal to identify positive sample pairs (i.e., the similarity between each sample and itself). Use the mask to select the elements on the diagonal in the similarity matrix, i.e., positive samples, and the elements on the off-diagonal, i.e., negative samples. Combine the similarity values ​​of positive and negative samples into a logits tensor, where the positive sample is located at the first position of each row and the rest are negative samples.

[0015] S2-4. Create a labels tensor of all zeros, because the target of the positive sample is always the first position, corresponding to the first element of each row in the logits tensor. Use the cross entropy loss to calculate the final contrast loss value, where logits is used as the predicted value and the labels tensor of all zeros is used as the target value;

[0016] S3. Use a variety of image enhancement methods to enhance the contrast of the image; this is to increase the diversity of image data and improve the generalization ability of the model.

[0017] Preferably, in step S1-2, for different tasks or different data sets, it is necessary to specify the field in Template 1 so that the LLM can generate prompt words that best suit the task.

[0018] Preferably, in step S1, not only task-specific prompt words are used, but also the prompt templates provided by the original and simplest CLIP are fully utilized to form prompt words; the prompt words of CLIP roughly locate the categories in the embedding space globally, while task-specific prompt words are more refined localizations that can distinguish the differences between similar categories; in step S1, two methods of weighting prompt words are used to jointly predict the categories of test samples, that is, the prediction vectors obtained from the prompt words of CLIP and the prediction vectors obtained from task-specific prompt words are weighted and added together to form a visual-text association prediction.

[0019] Preferably, in step S2, the contrastive loss trains the model by minimizing the similarity difference between positive samples and negative samples, so as to learn to distinguish the feature representations of different samples; the contrastive loss is added to the loss of the entire model to optimize the feature distribution in the adapter, so as to obtain more accurate predictions.

[0020] Preferably, in step S1, the visual features and the corresponding labels are respectively represented as and

[0021] F = VisualEncoder(X K )

[0022] L = OneHot(Y N )

[0023] where N represents N classes of training samples, K represents the number of N classes of training samples, X K represents N categories and K labeled images for each class, and Y N is the label.

[0024] Preferably, in step S2, the similarity matrix is:

[0025]

[0026] where τ is the temperature hyperparameter used to scale the similarity values;

[0027] The similarity between the positive samples and the negative samples is:

[0028] S + = S ⊙ M, S - = S ⊙ (1 - M),

[0029] where ⊙ represents element-wise multiplication and 1 is a matrix of all 1s;

[0030] The contrast loss is:

[0031]

[0032] Where H is the cross entropy (CE) loss function.

[0033] The beneficial effects of the present invention are:

[0034] 1. The CLIP model used in the present invention demonstrates a deep understanding of cross-modal information by virtue of its unique ability to process image and text pairs; through prompt optimization, the recognition performance of CLIP is significantly enhanced; these prompts are like navigation lighthouses, guiding the model to lock on key visual features in the image and provide necessary contextual information, thereby effectively reducing uncertainty and ambiguity in the recognition process. More importantly, the present invention strengthens the generalization ability of the model through rich descriptions of natural language, so that CLIP can handle fine-grained categories or complex concepts with ease. Therefore, the method of the present invention not only reduces the dependence on large-scale annotated data sets, but also improves the learning ability of the model in small or even zero sample situations, and enhances the adaptability and generalization ability of the model in the face of new environments; by simply and cleverly using task-specific prompts, it is possible to adjust and optimize the performance of CLIP without changing the model architecture or retraining, greatly simplifying the deployment process of the model, making it a powerful tool for solving a series of challenges in image recognition and description tasks.

[0035] 2. The present invention has shown significant advantages in the field of small sample learning by adopting a contrastive learning loss function and optimizing the feature representation strategy. First, the contrastive learning loss of the present invention enhances the discriminability of features, effectively solves the problem of unclear sample boundaries in small sample sets, and greatly reduces the probability of misclassification between similar samples, thereby significantly improving the accuracy of classification. Secondly, the present invention realizes dynamic optimization of feature representation. This strategy not only improves the model's adaptability to new samples, but also effectively avoids overfitting, ensuring that the model still has strong generalization capabilities when the amount of data is limited. The present invention simplifies the storage requirements of the model, optimizes computing efficiency, and provides a more efficient and robust solution for small sample learning. It is particularly suitable for scenarios where data is scarce, and shows significant effects in improving model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a framework diagram of the Tip-Adapter of the prior art;

[0037] Figure 2 It is a model framework diagram of a small sample learning method for task-adaptive association learning of the present invention;

[0038] Figure 3 It is a dimensionality reduction visualization diagram of the text features of the embodiments of the present invention;

[0039] Figure 4 It is a sample diagram in the PlantDoc and PlantVillage datasets of the embodiments of the present invention;

[0040] Figure 5 It is a comparison diagram of the sample prediction results of the embodiments of the present invention. Detailed implementation manners

[0041] Next, the related technologies in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] As Figures 1 - 5 shown, Figure 1 It is a framework diagram of Tip-Adapter (Zhang, Renrui et al. “Tip-Adapter: Training-free Adaption of CLIP for Few-shot Classification.” European Conference on Computer Vision (2022).).

[0043] Compared with the method of Figure 1 this embodiment has three aspects of improvement. The main improvement idea is to enhance the performance of the model by increasing visual relevance and multimodal relevance under the multimodal framework:

[0044] First, use task-specified prompts (TSP) to assist the original prompts.

[0045] The original CLIP-based prompt was: a photo of [CLS]. CLIP obtained the text embedding corresponding to the text category of the class by writing the class name into such a sentence and then feeding it into the text encoder. In the present invention, for different tasks and classes, task-specific prompts are designed. The specific operation is as follows: First, templates for the task are designed, such as "Utilize the knowledge of experts in this field to generate more accurate prompts for the class [CLS]." (Template 1) and "Use vivid and intuitive descriptive words to directly describe the visual features presented by this class." (Template 2). By designing these two templates, it can effectively avoid the generated prompts from being general and containing abstract adjectives, so that CLIP cannot accurately understand and locate to the appropriate position in the joint embedding space (the image and text feature embeddings learned by CLIP through contrastive learning is an aligned joint embedding space, that is, the text and images of the same class are aligned in the embedding space), and then find the most similar image features. Second, the above two templates are input into the Large Language Models (LLMs) in sequence. That is, first write the class name into Template 1 and input it into the LLM, and one or more prompt words will be obtained. Then, in a conversational manner, further input Template 2 into the LLM, and optimized prompt words will be obtained. By traversing all the classes in the dataset, a series of task-specific prompt words can be obtained and saved for use in subsequent model inference. It should be noted that for different tasks or different datasets, the field needs to be specified in Template 1 so that the LLM can generate the most task-appropriate prompt words.

[0046] In addition, the present invention not only utilizes task-specific prompt words but also fully utilizes the prompt templates provided by the original and most basic CLIP to form prompt words. The prompt words of CLIP globally locate the class to the approximate position in the embedding space, while the task-specific prompt words are more refined localizations that can distinguish the differences between similar classes. Therefore, the present invention adopts a method of weighting the two types of prompt words to jointly predict the class of the test sample, that is, the prediction vectors obtained by the prompt words of CLIP and the prediction vectors obtained by the task-specific prompt words are weighted and added together to form a visual-text association prediction.

[0047] II. Enhancing the discriminability of training features using contrastive learning

[0048] Since small sample learning is extremely dependent on very few training samples, the representation ability of the training samples is directly related to the prediction accuracy of the test data. Compared with the aforementioned literature that directly saves the image feature set of the training samples to query and compare with the features of the test samples, the present invention proposes to enhance the features of the feature set to increase the discriminability between the features. First, a linear layer network adapter is designed, the size of which is consistent with the transpose of the size of the feature set, and the network is initialized using the feature set extracted by the image feature extractor of CLIP. (This step is consistent with Tip-Adapter) Inputting the features of the input image into the adapter can obtain an affinity matrix affinity, which represents the similarity matrix between the input features and the training features. Normalizing affinity to obtain affinity_norm, taking this variable as the object of optimization, a contrastive learning loss is designed to narrow the distance between positive samples and push the distance between negative samples away, thereby learning useful feature representations. Secondly, the similarity matrix between the feature vectors (affinity) is calculated by multiplying the feature vector matrix with its transpose and dividing it by the temperature parameter temperature. The temperature parameter controls the scale of the resulting similarity, with lower temperatures making the differences in the similarity matrix more pronounced, while higher temperatures make these differences more gradual. Then create a Boolean mask matrix mask with True on the diagonal to identify positive sample pairs (i.e., the similarity of each sample to itself), and select the elements on the diagonal (positive samples) and the elements on the off-diagonal (negative samples) in the similarity matrix through the mask. Combine the similarity values ​​of the positive and negative samples into a logits tensor, where the positive samples are located in the first position of each row and the rest are negative samples. Create an all-zero labels tensor, because the target of the positive sample is always the first position, corresponding to the first element of each row in the logits tensor. Use the cross entropy loss to calculate the final contrastive loss value, with logits as the predicted value and the all-zero labels tensor as the target value. In general, this contrastive loss trains the model by minimizing the difference in similarity between positive and negative samples, thereby learning to distinguish the feature representations of different samples. This contrastive loss is added to the loss of the entire model to optimize the feature distribution in the adapter, so that more accurate predictions can be obtained. The network itself also has a classification loss, which is the prediction vector (Cache logits ) and the prediction vector of multimodal association (ω*CLIP logits +(1-ω)*TSP logits ) and the cross entropy loss between the predicted vector and the true label.

[0049] 3. Task-Specific Image Enhancement

[0050] For specific tasks, if visually similar samples are included, ordinary data augmentation methods cannot meet the requirements. For example, the default image augmentation in CLIP only includes two operations: scaling and cropping. Therefore, we propose to use diverse image augmentation methods. And we propose to only augment the training data, which, combined with the second improvement above, can effectively improve the discriminative representation ability of the training sample features. Enhance the contrast of the images. This is to increase the diversity of the image data and improve the generalization ability of the model.

[0051] The key points of this embodiment are as follows: 1. The task-specific prompt adjusts and optimizes the performance of CLIP in few-shot recognition tasks. 2. The training feature optimization supervised by contrastive learning reduces the probability of classification errors between similar samples. 3. The data augmentation technology can enrich the training samples.

[0052] Embodiment

[0053] This embodiment proposes an adaptive relationship learning method to solve the fine-grained few-shot classification problem in plant disease recognition, in order to learn the visual and language discrimination relationships between disease categories. The method framework of this embodiment is as Figure 2 shown. Generally speaking, this embodiment uses the frozen visual and text encoders of CLIP to extract the features of images and labels, and uses a learnable linear layer adapter to store and fine-tune the features of the few-shot training set. The prediction of this network consists of two parts: visual similarity mapping and cross-modal semantic alignment. The visual similarity mapping depends on the uniqueness of the training set features. Therefore, this embodiment introduces a contrastive method to learn more discriminative features. In terms of cross-modal semantic registration, this embodiment uses TSP, which supplements specific prior information for the plant disease prediction task to improve the accuracy of fine-grained classification. At the same time, this embodiment enhances the contrast of the training images to improve the robustness and generalization of visual similarity. Finally, this embodiment proposes a fusion prediction that combines the predictions of TSP and handcrafted prompts.

[0054] Visual similarity mapping

[0055] The methods in the prior art only store the image features in a memory bank for calculating the similarity between test samples. These methods focus on coarse-grained classification and ignore the inter-class feature similarity for fine-grained classification. This embodiment constructs a key-value cache model from the CLIP features extracted from the few-shot data to optimize the training features and improve the visual similarity mapping between the training samples and the test samples. Given a few-shot learning dataset that contains K training samples of N classes, which includes N classes and K labeled images for each class, denoted as X K , with the label Y N. We constructed a key-value cache model to adapt the training features to the CLIP model, which contains few-shot knowledge about N classes. For training samples, this embodiment uses CLIP to extract C-dimensional L2-normalized visual features and converts the true labels into one-hot vectors. The visual features and the corresponding labels are respectively denoted as and

[0056] F = VisualEncoder(X K )

[0057] L = OneHot(Y N )

[0058] In this embodiment, the labels are hidden in the cache model in Figure 2 . The visual features are regarded as keys of the key-value cache model, and the labels are regarded as values. The knowledge extracted from the few-shot learning dataset is stored in the cache model for subsequent fine-tuning.

[0059] To improve the discriminability of the cache model, this embodiment introduces a self-contrast method. First, we calculate the affinity between the training features and the test features. The C-dimensional visual features of the test images extracted by CLIP can be denoted as as a query to retrieve the most similar samples from the cache model. The affinity between the query (test features) and the keys (training features) can be calculated by matrix-vector multiplication:

[0060] A = exp(-β(1 - fF T )),

[0061] where and β represent modulation hyperparameters. In the case of a few samples, the amount of data is very limited, and it may be difficult to directly optimize the training features to capture the complex patterns and variations between the data. Instead, by optimizing the similarity (affinity) between the samples, the model can better utilize the contrast information between these limited samples and effectively utilize the supervision signals provided by each sample instance, thereby achieving more effective learning on a small number of samples.

[0062] Then, this embodiment constructs an affinity relationship contrast strategy. To bring the positive sample pairs closer and pull the negative sample pairs farther apart, this embodiment constructs a self-similarity matrix:

[0063]

[0064] where τ is the temperature hyperparameter used to scale the similarity values.

[0065] Then, in this embodiment, an indicator function is constructed for positive and negative samples. Positive sample pairs are composed of features of the same sample, usually defined by a mask \(M\in\{0,1\}\). N×N where \(M\) ii = 1 indicates diagonal elements (i.e., positive sample pairs), and the rest are zero. Negative sample pairs are all non - diagonal elements.

[0066] In this embodiment, the similarities of positive and negative samples are extracted respectively:

[0067] \(S\) + = \(S\odot M\), \(S\) - = \(S\odot(1 - M)\),

[0068] where \(\odot\) represents element - wise multiplication, and 1 is a matrix of all 1s. Combine the similarities of positive and negative samples:

[0069]

[0070] where \(B\) is the batch size, and set the label vector \(target\) i = [0, 0, 0, …, 0].[[]END]]

[0071] Finally, to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, the self - contrastive loss can be expressed as:

[0072]

[0073] where \(H\) is the cross - entropy (CE) loss function.

[0074] The logarithmic position in CE is the concatenation of the similarities of positive and negative samples. For each sample, the first element of its logarithmic vector is the similarity to itself (i.e., the positive sample), and the remaining elements are the similarities to other samples (i.e., the negative samples). The target label \(target\) is a vector of all zeros. This is because, for each sample, in this embodiment, it is desired that the first element of the logarithmic vector (i.e., the similarity to itself, i.e., the positive sample) is the maximum value, so that the similarity of each sample to itself is the maximum value and the similarity to other samples is the minimum value.

[0075] In addition, to enhance the model's perception of the visual features of leaves, this embodiment also proposes a customized data augmentation technique, namely Contrastive Enhancement Augmentation (CEA), to address two key challenges inherent in traditional visual encoders: one is its general design, which may not be able to fully utilize the unique features of plant leaves; the other is the robustness of the feature representation in the cache model, which is particularly important for the case of limited samples.

[0076] The theoretical basis behind CEA is that the standard visual coding architecture is not designed specifically for plant diseases, and there is a performance gap because it cannot detect complex details and variations in diseased leaves in the best way. By enhancing the contrast of leaf images, CEA can emphasize the visual boundaries between healthy and diseased tissues, thus enriching the discriminative information available to the model. Therefore, this enhancement can promote a more detailed understanding of disease symptoms and improve the model's ability to detect subtle disease indicators.

[0077] In addition, CEA is also crucial for enhancing feature robustness, especially in the case of a small number of samples. It can increase the contrast of training images, thereby enhancing feature diversity and maximizing the extraction of information from each sample, which is crucial for datasets with scarce data. This will improve the generalization ability, reduce the risk of overfitting, and improve the accuracy in a low-data environment.

[0078] Enhanced image X aug is obtained by applying a contrast enhancement function f to the pixel intensity values of the original image X (for illustration purposes) CEA defined as

[0079] X aug = f CEA (X) = α(X - μ) + μ,

[0080] where μ is the mean of X and α is a parameter that controls the degree of contrast enhancement.

[0081] Cross-modal semantic alignment

[0082] This embodiment adopts an innovative prompt design method of "Task Specific Prompts" (TSP), which aims to enrich text embeddings with a flexible semantic range by integrating fine-grained category information into the prompt construction process and improve cross-modal semantic alignment, thereby optimizing the performance of visual language models in few-shot learning tasks. Different from traditional handcrafted prompts, the TSP strategy is specifically designed to focus on and learn the important features of specific categories.

[0083] This embodiment specifically designs some templates for plant disease recognition. Researchers can use these templates in sequence, that is, use the first template to generate a prompt and then use the next template to optimize the generated prompt. The templates designed in this embodiment are as follows:

[0084] · Utilize the knowledge of domain experts to establish a more accurate prompt for [CLS].

[0085] · Directly describe the visual features of this plant disease with vivid and intuitive descriptions.

[0086] [CLS] represents the disease category to be described.

[0087] To address the challenges of the specific task of plant disease recognition, each prompt in ACIP is carefully designed around a specific plant disease to ensure the relevance and accuracy of the description. For example, for "apple scab leaves", the prompt used in this embodiment is "a photo of an apple scab leaf showing dark, rough scab-like lesion characteristics" to accurately convey the unique characteristics of apple scab. As for "tomato early blight leaves", this embodiment uses "a picture depicting early circular black spots on tomato leaves" to accurately focus on the early blight symptoms. Through such descriptions, TSP not only summarizes the general appearance of the category but also delves into the specific visual characteristics of the disease, such as color, shape, and distribution pattern, which are the keys to distinguishing different disease types.

[0088] In Figure 3 , this embodiment visualizes the text embeddings of all categories in the PlantDoc dataset under two prompts: CLIP handcrafted prompts and TSP prompts. After optimization by TSP, the text features become more discriminative, indicating that expanding category semantics can improve classification potential. This embodiment highlights the feature distributions of different levels of confusion before and after prompt optimization and provides the original confusing samples at the Figure 3 bottom. Pentagons represent highly confused categories. Under the handcrafted prompts, the features of "tomato early blight leaves" and "tomato late blight leaves" almost overlap, but the specific disease descriptions will significantly increase the distance. Similarly, for the moderately confused "tomato moldy leaves" and "tomato leaves", the improved prompts expand the feature distance by about 8 times. Even for samples with a low level of confusion (mild), TSP can further increase the inter-class distance. These results show that the prompt optimization TSP in this embodiment can effectively enhance the discriminability of text features and improve the model's ability to distinguish visually similar plant diseases.

[0089] TSP directly maps complex disease features in the image that are difficult to intuitively identify by pixels through language, which prompts the model to go beyond the pixel level when understanding the image and directly reach the essential attributes of the disease. In this way, TSP not only enriches the semantic information received by the model but also promotes the close combination between visual features and highly relevant language descriptions.

[0090] Fusion prediction of multimodal models

[0091] For the test images, this embodiment uses CLIP to extract C-dimensional visual and text features, denoted as and In addition, this embodiment sets category prompts to obtain TSP-guided text embeddings Then, the calculation formula for prediction through visual-text similarity is

[0092]

[0093] In addition, the small-sample knowledge retrieved from the cache model also helps predict visual-visual similarity,

[0094]

[0095] wherein

[0096] Finally, this embodiment integrates the two parts of the prediction to obtain a fused prediction:

[0097] logits = logits text + γlogits visual

[0098] where γ represents the residual ratio. The final predicted category is the category corresponding to the maximum value of logits.

[0099] This embodiment conducts experiments on two widely used plant disease recognition datasets: Figure 4 Samples of these two datasets are shown. The PlantDoc dataset was collected under natural field conditions and contains 2,569 images of 27 categories (17 disease categories and 10 healthy categories) of 13 crops. In contrast, the widely used PlantVillage dataset was launched by Pennsylvania State University for plant disease diagnosis and contains 54,305 healthy and diseased leaf images of 14 crop varieties, divided into 38 categories (26 diseases and 12 healthy). However, an important limitation of the PlantVillage database is that these images were not taken in a natural environment but were collected in a laboratory environment with a simple background. As Figure 5 shown, PlantDoc consists of images of diseased plants that have different environments, multiple leaves, fruits, and different lighting and illumination conditions. The backgrounds of PlantVillage images are clear, and the capture and illumination conditions are uniform. Figure 5 Samples in the PlantDoc and PlantVillage datasets. The PlantDoc dataset was taken from a natural environment and is more challenging due to the presence of various backgrounds and leaves. The PlantVillage dataset was collected in a laboratory environment with a simple background, and there is only one leaf in one image.

[0100] Experimental setup

[0101] In the aspect of few-shot learning, this embodiment compares the performance of training sets with 1, 2, 4, 8, and 16 samples per class and tests on the complete test set. This embodiment selects ResNet-50 as the visual encoder of the CLIP backbone and selects Transformer as the text encoder. This embodiment uses a pre-trained CLIP model and freezes the parameters during training. The image augmentation is consistent with the preprocessing protocol of CLIP, which includes random cropping, resizing, and random horizontal flipping. The contrast enhancement parameter α is set to 0.5. The parameter search function automatically learns β. The weighted parameter ω is set to 0.5. For the temperature parameter τ and the weight parameter γ of the ratio learning loss, finally τ is selected as 0.5 and γ is selected as 0.5. The learnable cache model is a linear layer, whose input and output sizes are used as training features (keys) and are initialized by the transpose of the keys. In the training phase and the test phase, this layer is fine-tuned using a batch size of 64 and 256, a learning rate of 0.001, and the AdamW optimizer with a cosine scheduler. For most results, this embodiment sets the training epochs to 40.

[0102] According to whether the cache model is learnable, this embodiment compares two methods: training-free (CALP) and training-required (CALP-F, CALP-Finetune). For the former, the visual-visual similarity is directly calculated from the affinity, while for the latter, the cache model is fine-tuned. This embodiment also compares the results of predicting only using TSP (referred to as "new prompt") and using fusion prediction.

[0103] This embodiment conducts a performance comparison of zero-shot CLIP, Tip-Adapter, and Tip-Adapter-F (Tip-Adapter with fine-tuning). Zero-shot CLIP uses pre-trained knowledge for classification and does not require additional training. Tip-Adapter stores training features and is used for prediction. Tip-Adapter-F converts training features into learnable features. The classification results (classification accuracy, expressed as a percentage) of the PlantDoc dataset are shown in Table 1:

[0104] Method 1 2 4 8 16 24 36 48 64 128 Zero-shot CLIP 11.46 11.46 11.46 11.46 11.46 11.46 11.46 11.46 11.46 11.46 Zero-shot CLIP + New Prompt 19.90 19.90 19.90 19.90 19.90 19.90 19.90 19.90 19.90 19.90 Zero-shot CLIP + Fusion Prediction 17.23 17.23 17.23 17.23 17.23 17.23 17.23 17.23 17.23 17.23 Tip-Adapter 13.89 20.37 31.73 42.04 55.19 59.62 59.14 61.73 59.37 59.42 CALP + New Prompt 25.01 32.08 39.77 47.84 55.21 59.09 57.02 59.75 58.22 57.37 CALP + Fusion Prediction 23.92 37.78 54.23 56.27 56.90 60.25 59.05 61.18 59.96 58.91 Tip-Adapter-F 34.41 44.72 60.08 70.30 77.84 80.48 83.74 85.64 86.90 90.84 CALP-F + New Prompt 28.48 38.96 45.77 66.38 75.23 82.04 84.40 85.71 87.49 90.82 CALP-F + Fusion Prediction 37.80 54.06 65.05 75.72 79.17 82.38 84.66 86.28 87.66 91.66

[0105] The results of the PlantVillage dataset are shown in Table 2:

[0106] Method 1 2 4 8 16 24 36 48 64 128 Zero-shot CLIP 24.27 24.27 24.27 24.27 24.27 24.27 24.27 24.27 24.27 24.27 Zero-shot CLIP + New Prompt 28.93 28.93 28.93 28.93 28.93 28.93 28.93 28.93 28.93 28.93 Zero-shot CLIP + Fusion Prediction 28.35 28.35 28.35 28.35 28.35 28.35 28.35 28.35 28.35 28.35 Tip-Adapter 26.21 29.71 30.29 40.58 41.75 43.88 44.66 44.75 45.05 48.74 CALP + New Prompt 29.71 34.76 33.40 42.91 42.33 47.38 45.24 44.36 46.80 46.41 CALP + Fusion Prediction 28.35 34.95 32.62 43.11 44.66 47.18 45.44 44.75 46.80 47.38 Tip-Adapter-F 37.48 36.89 38.45 50.29 55.34 61.75 61.17 58.83 63.50 64.08 CALP-F + New Prompt 38.25 40.97 40.19 46.60 57.28 60.58 62.14 59.42 63.69 61.94 CALP-F + Fusion Prediction 42.14 42.52 42.33 53.20 60.19 63.11 63.11 61.55 64.47 65.63

[0107] Table 2

[0108] Partial classification results: Figure 5Sample prediction result comparison for three methods: Tip-Adapter, CALP prompt, and CALP fusion. The dark black bars are correct predictions, the light black-red bars are incorrect predictions, and the gray bars are other predictions. The labels of the top 3 predictions are shown.

[0109] In summary, the present invention strengthens the generalization ability of the model through rich descriptions in natural language, enabling CLIP to handle even fine-grained categories or complex concepts with ease. The method of the present invention not only reduces the dependence on large-scale labeled datasets, but also improves the learning ability of the model in the few-shot or zero-shot scenarios, enhances the adaptability and generalization ability of the model in the face of new environments. The present invention greatly simplifies the model deployment process, making it a powerful tool for solving a series of challenges in image recognition and description tasks.

[0110] It should be emphasized that the above are only the preferred embodiments of the present invention, and do not constitute any form of limitation to the present invention. Any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A small sample learning method for task-adaptive association learning, characterized in that: The method comprises the following steps: S1. Use task-specific prompts to supplement existing prompts; specifically: S1-1. Design a template for the task, the template comprising: Template 1 and Template 2. Template 1 comprises: utilizing the knowledge of experts in the field to generate more accurate prompts for the category; Template 2 comprises: using figurative and intuitive descriptive words to directly describe the visual features presented by the category; S1-2, input the two templates designed in step S1-1 into the large language model in sequence, that is, first write the category name into template one and input it into LLM, you will get one or more prompt words, and then further input template two into LLM in a dialogue way, you will get optimized prompt words; traverse all categories in the data set to get a series of task-specific prompt words, save them, and use them in subsequent model reasoning; S2. Use contrastive learning to enhance the discriminability of training features; specifically including: S2-1. Design a linear layer network adapter, the size of which is consistent with the transpose of the size of the feature set, and initialize the network using the feature set extracted by the CLIP image feature extractor; input the features of the input image into the adapter to obtain an affinity matrix affinity, which represents the similarity matrix between the input features and the training features; normalize affinity to obtain affinity_norm, and use this variable as the object of optimization; S2-2, calculating the similarity matrix between the eigenvectors by multiplying the eigenvector matrix by its transpose and dividing by the temperature parameter; S2-3. Create a Boolean mask matrix mask with True on the diagonal. Use the mask to select the elements on the diagonal of the similarity matrix, i.e., positive samples, and the elements on the off-diagonal, i.e., negative samples. Combine the similarity values ​​of the positive and negative samples into a logits tensor, where the positive samples are located at the first position of each row and the rest are negative samples. S2-4. Create a zero labels tensor and use the cross entropy loss to calculate the final contrast loss value, where logits are used as the predicted value and the zero labels tensor is used as the target value. S3. Use a variety of image enhancement methods to enhance the contrast of the image.

2. The small sample learning method for task-adaptive association learning according to claim 1, characterized in that: In step S1-2, for different tasks or different data sets, the field needs to be specified in template 1 so that LLM can generate prompt words that best fit the task.

3. The small sample learning method for task-adaptive association learning according to claim 1, characterized in that: In the step S1, not only the task-specific prompt words are used, but also the prompt template provided by the original and simplest CLIP is fully utilized to form the prompt words; the prompt words of CLIP globally locate the category to the approximate position of the embedding space, while the task-specific prompt words are more precise local positioning; the step S1 adopts a weighted method of two prompt words to jointly predict the category of the test sample, that is, the prediction vector obtained by the prompt words of CLIP and the prediction vector obtained by the task-specific prompt words are weighted and added to form a prediction of visual-text association.

4. The small sample learning method for task-adaptive association learning according to claim 1, characterized in that: In step S2, the contrast loss trains the model by minimizing the similarity difference between positive samples and negative samples, thereby learning to distinguish feature representations of different samples; the contrast loss is added to the loss of the entire model to optimize the feature distribution in the adapter, thereby obtaining a more accurate prediction.

5. The small sample learning method for task-adaptive association learning according to claim 1, characterized in that: In step S1, the visual features and the corresponding labels are respectively represented as and F=VisualEncoder(X K ) L=OneHot(Y N ) Among them, N represents N-type training samples, K represents the number of N-type training samples, and X K Indicates that there are N categories and K labeled images per category, Y N is the label, and C represents the dimension of the visual feature.

6. The small sample learning method for task-adaptive association learning according to claim 1, characterized in that: In step S2, the similarity matrix is: Where τ is a temperature hyperparameter used to scale the similarity value; A = exp(-β(1-fF T )),in and β represent modulation hyperparameters, f represents C-dimensional visual features, and F represents visual features; Positive sample pairs consist of features of the same sample, defined by a mask M∈{0,1} N×N Among them, M ii =1 represents the diagonal elements, and the rest are zero; negative sample pairs are all non-diagonal elements; The similarity between the positive sample and the negative sample is: S + =S⊙M,S - =S⊙(1-M), Where ⊙ represents element-wise multiplication, 1 is a matrix of all 1s; S + , S - Represent the similarity of positive samples and negative samples respectively; merge the similarity of positive and negative samples: Where B is the batch size and the label vector target is set i =[0, 0, 0, ..., 0]; The contrast loss is: Where H is the cross entropy CE loss function.

Citation Information

Patent Citations

  • Visual language understanding task processing method and system

    CN116432026A

  • Small sample image classification optimization method based on supervised comparative learning and multi-task setting

    CN116524242A