Small sample class incremental learning method and device, electronic equipment and readable storage medium

By integrating the fine-grained expert model and multimodal feature fusion based on CLIP, the problem of poor CLIP performance in the FG-FSCIL scenario is solved, and more efficient fine-grained category recognition and new category learning are achieved.

CN120339635APending Publication Date: 2025-07-18HUNAN FIRST NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410284957.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the Fine Grained Small Sample Class Incremental Learning (FG-FSCIL) scenario, the CLIP-based learning method has poor performance, making it difficult to effectively identify the differences between fine Grained categories, and it is difficult to confuse old and new categories and adapt to new categories.

Method used

The image integration encoder is used to combine the fine-grained expert model and the CLIP text encoder to generate feature prototypes by fusing image features and text features, and to enhance the generalization ability of the model using supervised contrast learning and virtual category data.

Benefits of technology

The performance of the model in the FG-FSCIL scenario is improved, the ability to distinguish fine-grained categories is enhanced, the confusion between new and old categories is reduced, and the learning effect of new categories is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339635A_ABST
    Figure CN120339635A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample class incremental learning method, and particularly relates to a small sample class incremental learning method and device, electronic equipment and a readable storage medium. The method comprises the steps that a to-be-processed image is input into an image integrated encoder for encoding processing, target image features are obtained, the image integrated encoder comprises a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used for recognizing the difference between fine-grained categories; inputting the to-be-processed text into a CLIP text encoder for encoding processing to obtain a target text feature; and determining a learning result based on the target text feature and the target image feature. The method provided by the invention has relatively good performance in an FG-FSCIL scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to a few-shot class-incremental learning method, and specifically relates to a few-shot class-incremental learning method, device, electronic device and readable storage medium. Background Art

[0002] Traditional deep learning (DL) models have achieved great success in computer vision tasks, mainly due to massive training samples and a large amount of computing resources. However, in the face of a dynamic world, these models often need to retrain on sufficient samples for new and old classes to adapt to new classes while retaining old knowledge. In dynamic scenarios with limited data, the storage of old classes, the annotation of new classes, and the demand for a large amount of computing resources make this strategy costly and less feasible.

[0003] Few-shot class-incremental learning (FSCIL) focuses on continuously learning new classes with limited samples while retaining old knowledge. Comparing the application of contrastive language-image pre-training (CLIP) methods or models in FSCIL, it mainly uses their significant generalization ability to address catastrophic forgetting.

[0004] However, in the fine-grain few-shot class-incremental learning (FG-FSCIL) scenario, due to the high similarity between classes, the FG-FSCIL scenario has the following characteristics:

[0005] 1) It is difficult to distinguish base classes. These fine-grained classes usually have high visual similarity. Even during basic training, it is already a challenge for general models to learn these subtle differences between fine-grained classes, which will pose a challenge for general models to identify old classes during the basic session.

[0006] 2) The dilemma of confusing base classes and new classes. In addition to identifying old classes, these models also need to continuously learn new classes. The introduced new classes may be highly visually similar to some old classes. This not only increases the difficulty of classifying old classes but also poses a challenge for accurately classifying new classes, resulting in confusion between old and new classes.

[0007] 3) The dilemma of adapting to new classes. In addition to the introduced new classes being similar to old classes, the available samples of these new classes are very limited, which makes it difficult for models to learn fine-grained discriminative information about new fine-grained classes. Therefore, general models face challenges in adapting to these new classes.

[0008] Given the typical characteristics of the above FG-FSCIL scenario, the performance of the CLIP-based learning method in the FG-FSCIL scenario is poor. Summary of the Invention

[0009] The technical problem to be solved by the present invention is that the performance of the CLIP-based learning method in the FG-FSCIL scenario is poor. To solve the above problem, the present invention provides a few-shot class incremental learning method, device, electronic device and readable storage medium.

[0010] The content of the present invention includes:

[0011] In a first aspect, an embodiment of the present invention provides a few-shot class incremental learning method, including:

[0012] Input the image to be processed into an image integration encoder for encoding to obtain target image features. The image integration encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained categories;

[0013] Input the text to be processed into a CLIP text encoder for encoding to obtain target text features;

[0014] Determine the learning result based on the target text features and the target image features.

[0015] Optionally, the step of inputting the image to be processed into an image integration encoder for encoding to obtain target image features includes:

[0016] Input the image to be processed into the CLIP image encoder to obtain a first image feature, and input the image to be processed into the fine-grained expert model to obtain a second image feature;

[0017] Fuse the first image feature and the second image feature to obtain the target image feature.

[0018] Optionally, the step of fusing the first image feature and the second image feature to obtain the target image feature includes:

[0019] Input the first image feature into a first fully connected layer for dimension adjustment to obtain a third image feature, and the dimension of the third image feature is the same as the dimension of the second image feature;

[0020] Average the third image feature and the second image feature to obtain the target image feature.

[0021] Optionally, the method further includes:

[0022] Iteratively train the integrated network to be trained based on the training dataset to obtain an integrated network, where the integrated network includes the image integrated encoder and the CLIP text encoder;

[0023] Among them, the training dataset includes old-class data and virtual-class data, and the virtual-class data is obtained by augmenting the old-class data.

[0024] Optionally, before iteratively training the integrated network to be trained based on the training dataset to obtain an integrated network, the method further includes:

[0025] Based on the target image features corresponding to the target class and the text features corresponding to the target class, determine the feature prototype of the target class.

[0026] Optionally, before obtaining the feature prototype of the target class based on the target image features corresponding to the target class and the text features corresponding to the target class, the method further includes:

[0027] Obtain supplementary text information corresponding to the target class, where the supplementary text information is used to describe the target class;

[0028] Input the supplementary text information into the CLIP text encoder for encoding processing to obtain the text features corresponding to the target class.

[0029] Optionally, the fine-grained expert model is a ResNet18 with a self-attention mechanism.

[0030] In a second aspect, an embodiment of the present invention provides a few-shot class incremental learning device, including:

[0031] A first processing module, configured to input a to-be-processed image into an image integrated encoder for encoding processing to obtain target image features, where the image integrated encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained classes;

[0032] A second processing module, configured to input a to-be-processed text into the CLIP text encoder for encoding processing to obtain target text features;

[0033] A first determination module, configured to determine a learning result based on the target text features and the target image features.

[0034] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the few-shot class incremental learning method as described in the first aspect.

[0035] In a fourth aspect, an embodiment of the present invention provides a readable storage medium for storing a program, which when executed by a processor implements the steps in the few-shot class incremental learning method as described in the first aspect.

[0036] In an embodiment of the present invention, an image to be processed is input into an image integration encoder for encoding to obtain target image features. The image integration encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained classes; the text to be processed is input into a CLIP text encoder for encoding to obtain target text features; a learning result is determined based on the target text features and the target image features. Since a fine-grained expert model is integrated on the basis of the CLIP image encoder, and the fine-grained expert model can effectively learn the differences between fine-grained classes. Through the above settings, the ability of the image integration encoder is enhanced, enabling it to extract features with higher discriminability. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG Figure 1 is a flowchart of the few-shot class incremental learning method provided by an embodiment of the present invention;

[0038] FIG Figure 2 is a training schematic diagram of the image integration encoder and the CLIP text encoder provided by an embodiment of the present invention;

[0039] FIG Figure 3 is a schematic diagram of generating a Query view and a Key view provided by an embodiment of the present invention;

[0040] FIG Figure 4 is a schematic diagram of generating a feature prototype for each category provided by an embodiment of the present invention;

[0041] FIG Figure 5 is a schematic diagram of the integrated network structure provided by an embodiment of the present invention;

[0042] FIG Figure 6 is a schematic diagram of the performance comparison result provided by an embodiment of the present invention;

[0043] FIG Figure 7 is a schematic diagram of the few-shot class incremental learning device provided by an embodiment of the present invention;

[0044] FIG Figure 8 is a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In the embodiments of the present application, the term "and / or" describes the relationship between associated objects and represents three possible relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates an "or" relationship between the associated objects before and after. In the embodiments of the present application, the term "plural" refers to two or more, and other quantifiers are similar.

[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0047] Please refer to Figures 1 - 5 , Figure 1 which is a schematic flowchart of the small-sample class incremental learning method provided by the embodiments of the present invention. The method provided by the embodiments of the present invention can be used in the FG-FSCIL scenario to achieve fine-grained small-sample class incremental learning.

[0048] As Figure 1 shown, the method specifically includes the following steps:

[0049] Step 101: Input the image to be processed into the image integration encoder for encoding to obtain the target image features. The image integration encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained categories.

[0050] Step 102: Input the text to be processed into the CLIP text encoder for encoding to obtain the target text features.

[0051] Step 103: Determine the learning result based on the target text features and the target image features.

[0052] The embodiments of the present invention provide an integrated network, which can also be called an integrated network of arbitrary granularity based on CLIP. The integrated network includes an image integration encoder and a CLIP text encoder. Among them, the image integration encoder is obtained by integrating a CLIP image encoder and a fine-grained expert model. The text to be processed can also be called the prompt text.

[0053] For convenience of description, in this article, the integrated network of arbitrary granularity based on CLIP is called the AGEN network, the image integration encoder is denoted as f I , the CLIP image encoder is denoted as f I c , and the fine-grained expert model is denoted as f e, the CLIP text encoder is denoted as f T .

[0054] The fine-grained expert model is used to identify the differences between fine-grained categories, and its specific structure is not limited herein. Exemplarily, as an alternative implementation, the fine-grained expert model is a residual network. As another alternative implementation, the fine-grained expert model is ResNet18 with a self-attention mechanism.

[0055] In this embodiment, by integrating ResNet18 with a self-attention mechanism and the CLIP image encoder, the attention mechanism can be utilized to guide the model to pay more attention to the subtle differences between fine-grained categories. By the collaborative work of ResNet18 with a self-attention mechanism and the CLIP image encoder, the accuracy of the finally obtained learning result can be improved.

[0056] It should be understood that the specific structures of the CLIP text encoder, the CLIP image encoder, and ResNet18 with a self-attention mechanism can be referred to the descriptions in the related technologies and will not be elaborated herein.

[0057] In this embodiment, the execution order of step 101 and step 102 is not limited herein. As an alternative implementation, the text to be processed can be input into the CLIP text encoder for encoding while the image to be processed is input into the image integration encoder for encoding. As another alternative implementation, the image to be processed can be first input into the image integration encoder for encoding, and then the text to be processed is input into the CLIP text encoder for encoding.

[0058] In the embodiments of the present application, the image integration encoder includes a CLIP image encoder and a fine-grained expert model. The fine-grained expert model is used to identify the differences between fine-grained categories. Since the fine-grained expert model is integrated on the basis of the CLIP image encoder, and the fine-grained expert model can effectively learn the differences between fine-grained categories. Through the above settings, the ability of the image integration encoder is enhanced, enabling it to extract more discriminative features.

[0059] Optionally, in some embodiments, step 101 includes:

[0060] Input the image to be processed into the CLIP image encoder to obtain a first image feature, and input the image to be processed into the fine-grained expert model to obtain a second image feature;

[0061] Fuse the first image feature and the second image feature to obtain the target image feature.

[0062] Please refer to Figure 2, in this embodiment, the to-be-processed image is input into the CLIP image encoder and the fine-grained expert model respectively to obtain a first image feature and a second image feature. By fusing the first image feature and the second image feature, the final target image feature is obtained, so that the finally obtained target image feature has both distinctiveness and generalization ability.

[0063] There is no limitation on the specific way of fusing the first image feature and the second image feature to obtain the target image feature. Optionally, as an alternative embodiment, the first image feature and the second image feature are subjected to weighted summation processing to obtain the target image feature.

[0064] Optionally, as another alternative embodiment, the fusing of the first image feature and the second image feature to obtain the target image feature includes:

[0065] The first image feature is input into a first fully connected layer for dimension adjustment to obtain a third image feature, and the dimension of the third image feature is the same as that of the second image feature;

[0066] The third image feature and the second image feature are averaged to obtain the target image feature.

[0067] The feature dimension extracted by the CLIP image encoder is not consistent with the feature dimension extracted by the fine-grained expert model. By adjusting the dimension of the feature through the first fully connected layer, it is convenient to fuse the first image feature and the second image feature.

[0068] Specifically, in this embodiment, the third image feature and the second image feature are averaged. The feature obtained after the averaging process has both distinctiveness and generalization ability.

[0069] Optionally, in some embodiments, the method further includes:

[0070] Iteratively training the to-be-trained integrated network based on a training data set to obtain an integrated network, where the integrated network includes the image integrated encoder and the CLIP text encoder;

[0071] Among them, the training data set includes old-class data and virtual-class data, and the virtual-class data is obtained by performing an augmentation operation on the old-class data.

[0072] Please refer to Figures 2 - 5 , in this embodiment, the to-be-trained integrated network is trained on a basic session. During this process, the CLIP image encoder in the to-be-trained integrated network remains frozen. The loss function for training the image integrated encoder is the supervised contrast loss Lsup and the cross - entropy loss L ce . Supervised contrastive learning aims to minimize the distribution between samples of the same class while maximizing the separation between samples belonging to different classes. If the attention mechanism can guide the model to focus on a region that can achieve this goal, then the model can demonstrate strong capabilities in accurately classifying fine - grained classes. The supervised contrastive loss helps the image encoder learn the differences between fine - grained classes. In Figure 2 , the cross - entropy loss functions in the middle and at the bottom are used for classification and alignment with the text encoder respectively.

[0073] For supervised contrastive learning, given a batch of image - label pairs \((x, y)=\{x i , y\}\ b i=0 , the query view \(x q =\text{Aug}(x)\) and the key view \(x k =\text{Aug} k (x)\) will be generated by random augmentation. Then \(x q \) and \(x k \) will be input into \(g()\) to obtain the L2 - normalized tables \(q\) and \(k\), where \(g\) consists of an \(f I \) and a projection layer \(h\). \(q\) and \(k\) will be used to calculate the supervised contrastive loss. For this purpose, given an image \(x i \in x\), the supervised contrastive loss is defined as follows:

[0074]

[0075] where \(k + \) represents the set of positive samples, that is, the set of samples in \(k\) that belong to the same class as \(x i \), \) is the temperature parameter. Supervised contrastive learning mainly improves the performance of the model in the base session, but performs poorly in the incremental session. Supervised contrastive learning may cause the model to over - focus on the classification of old classes and ignore the prediction of new classes that appear in the future. Therefore, in this embodiment, virtual class data is generated by performing an augmentation operation \(\text{Aug}\) on the old - class data, enabling the model to imagine future new classes, thereby enhancing the generalization ability. For the cross - entropy loss \(L1\) used for classification, its definition is as follows:

[0076]

[0077] As Figure 3 shown, Rotational Aug is used to create virtual classes, while \(Q\) and \(K\) are random augmentations to generate supervised contrastive image data.

[0078] In this embodiment, the supervised contrastive loss helps to learn the differences between fine-grained categories, and the cross-entropy loss function is used for classification and alignment with the text encoder respectively. Meanwhile, the generalization ability of the model can be enhanced by augmenting the old category data with the Aug operation to generate virtual category data.

[0079] Optionally, in some embodiments, before iteratively training the to-be-trained integrated network based on the training dataset to obtain the integrated network, the method further includes:

[0080] Based on the target image features corresponding to the target category and the text features corresponding to the target category, obtain the feature prototype of the target category.

[0081] The features representing the category, that is, the feature prototype. In the related art, only text features are used as the feature prototype, and only class name information is used in the CLIP text information, which results in the lack of discriminability of the text features extracted by the CLIP text encoder. In the embodiments of the present application, the target image features extracted by the image integrated encoder show stronger discriminability. Therefore, in this embodiment, the average category image feature A I ={A I i} n i=1 is used as the feature prototype, where the subscript i represents the category and the superscript I represents the image feature.

[0082] Since the number of samples of the new category is limited, it is difficult to adapt to the new category by only using text features or image features to represent the category. In this embodiment, by fusing text features and image features to generate the feature prototype, the feature prototype can better represent the category, which can significantly enhance the discriminability of the feature prototype of each category, and thus improve the overall performance of the integrated network model.

[0083] In the incremental session, the number of samples of the new category is limited, and they may show a high degree of similarity to the old category. Therefore, it is also difficult to improve the discriminability of the feature prototype only relying on image features.

[0084] As Figure 4 shown, the fusion of multi-modal data can make the feature prototype show stronger distinguishing characteristics. During the test process, the virtual category and the corresponding real category belong to the same category, so they are fused with the same text features. As Figure 5 shown, in order to enhance the classification ability of fine-grained categories, ResNet18 with self-attention is integrated with the CLIP image encoder. After the CLIP text encoder and the image integrated encoder, a fully connected layer is added to adjust the dimension of the features for feature fusion.

[0085] Optionally, in some embodiments, before obtaining the feature prototype of the target category based on the target image features corresponding to the target category and the text features corresponding to the target category, the method further includes:

[0086] Obtaining supplementary text information corresponding to the target category, where the supplementary text information is used to describe the target category;

[0087] Inputting the supplementary text information into the CLIP text encoder for encoding processing to obtain the text features corresponding to the target category.

[0088] In this embodiment, supplementary text information is collected for each category from the dataset. The supplementary text information can describe the target category in more detail, achieving the effect of enhancing the text description, thereby generating more discriminative text features, and further improving the discriminability and distinctiveness of the feature prototype.

[0089] Exemplarily, as a specific embodiment, the specific manner of obtaining the feature prototype of the target category based on the target image features corresponding to the target category and the text features corresponding to the target category is as follows:

[0090]

[0091] where P i and F T i are the feature prototype of category i and the text features corresponding to category i respectively, and α is a hyperparameter.

[0092] It should be understood that, in some embodiments, these feature prototypes are composed of the original category and the virtual category. Since the virtual category and its corresponding original category are regarded as the same category during the test, the text features fused with the virtual category are the same as those fused with the original category.

[0093] In some embodiments, in the case of dimensional mismatch between the target text features and the target image features, a second fully connected layer is added after the CLIP text encoder to adjust the dimensions of the features output by the CLIP text encoder. As an alternative implementation, the fully connected layers after the CLIP image encoder and the CLIP text encoder are shared.

[0094] The text feature F T only matches the features extracted by the CLIP image encoder. However, the A I used for fusion is extracted by the image integration encoder. In some embodiments, the cross-entropy loss function L2 is used in the basic training to make the features extracted by the image integration encoder match the text features to a certain extent.

[0095] Specifically, the definition of L2 is as follows:

[0096]

[0097] Among them, F T base is the text feature of a set of base classes. Therefore, the total loss function L for training f I is as follows:

[0098]

[0099] Among them, β is a hyperparameter used to weigh the importance of L2.

[0100] Next, a specific dataset is used to verify the method provided in this application.

[0101] Aircraft and CUB200 are datasets widely used in the Fine-Grained Visual Categorization (FGVC) task. FGVC is dedicated to classifying subordinate categories belonging to the same superclass (such as specific types of birds), and the differences between these subordinate categories are very subtle.

[0102] Although the FSCIL benchmark dataset includes CUB200, a single fine-grained dataset is not sufficient to comprehensively evaluate the performance of the integrated network model provided in this application in the FG-FSCIL scenario. Therefore, Aircraft is incorporated into the FG-FSCIL scenario in this embodiment. Over time, new aircraft models will be designed, and it is very difficult to obtain labeled fine-grained data for these new models. Considering these factors, aircraft classification meets the requirements of the FG-FSCIL scenario.

[0103] Aircraft includes 100 different aircraft variants, which is the same as the number of categories in the miniImageNet dataset in the FSCIL benchmark dataset. Therefore, 60 basic aircraft categories are established in the miniImageNet category incremental format, and the remaining 40 aircraft categories are designated as new categories. The incremental learning of these 40 new categories will be in the form of 5-way 5-shot.

[0104] Thirty training samples and thirty test samples are selected for each category. In the incremental session, only 5 training samples are available for the new categories. Through the above method, a total of 6000 fine-grained aircraft images are collected. In addition, during the training phase, the size of these images will be adjusted to 256×256 and then randomly cropped to 224×224. Only 60 basic categories are tested in the basic phase. In the incremental phase, both the base classes and the new classes are tested. Only 60 basic categories are tested in the basic phase.

[0105] CUB200 is a fine-grained dataset consisting of 11,788 images, representing 200 different bird species. Among them, 5,994 images are used as training data, and the remaining 5,794 images are for testing. During the training phase, the size of each image is adjusted to 256×256 and then cropped to 224×224. The first 100 categories are designated as old categories, while the subsequent 100 categories are for incremental learning in the form of 10-way 10-shot.

[0106] MiniImageNet is a subset of the ImageNet-1k dataset, which contains 100 coarse-grained categories, with 600 images in each category, and each image has a size of 84×84. Among them, 60 categories are regarded as base classes, and the remaining 40 categories are designated as new classes. The structural format of the data for these new categories is 5-way 5-shot.

[0107] CIFAR100 is a coarse-grained dataset containing 100 categories, with 600 samples in each category, and each sample has a size of 32×32. Among them, 500 are for training and 100 are for testing. Among the total 100 categories, 60 are classified as base classes, and the remaining 40 are designated as new classes. Each incremental session introduces 5 new categories, and only 5 samples are available for each new category.

[0108] The CLIP model used in this embodiment is ViT-L / 14@336px. The encoder of the fine-grained expert model in this embodiment adopts ResNet18 backbone for miniImageNet, Aircraft100 and CUB200, and ResNet20 backbone for CIFAR100. In this embodiment, SGD with a momentum of 0.9 is used to optimize the model. In the base session, the initial learning rate for CIFAR100 and miniImageNet is 0.1, and the initial learning rate for CUB200 and Aircraft100 is 0.005. For the augmentation operation Aug, only one rotation is used in this embodiment to generate virtual categories. α and β for all datasets are set to 0.8 and 0.2 respectively to fuse text features and weigh the importance of L2. All experiments are conducted on an Nvidia Tesla A100 GPU.

[0109] In this paper, the method proposed in this application (abbreviated as AGEN) is compared with the methods in some related technologies on three FSCIL benchmark datasets and the Aircraft100 dataset. These methods include CEC, FACT, BiDist, ALICE, SSFE and SAVC (for specific details, please refer to the description of related technologies and will not be elaborated here).

[0110] For SSFE, it is only compared on the fine-grained dataset because it is an ultra-fine-grained few-shot class incremental learning method. Therefore, we added a comparison method GKEAL on the coarse-grained dataset. In addition, the performance of the original CLIP model is also shown. The detailed numerical values of CUB200 and Aircraft100 are shown in Table 2 and Table 1 respectively, and the performance curves on the miniImageNet and CIFAR100 datasets are as Figure 6 shown.

[0111] It should be understood that Table 1: Comparison with SOTA methods on Aircraft100, and the results of all other methods are implemented on the Aircraft100 dataset using the officially released code. Table 2: Comparison with SOTA methods on CUB20. After all, the "*" in the table indicates the results implemented using the officially released code

[0112] Table 1

[0113]

[0114] Table 2

[0115]

[0116] As can be seen from Table 2 and Table 1, CLIP performs poorly on the fine-grained dataset, while most FSCIL methods are better than it. One of the reasons for CLIP's poor performance in FG-FSCIL is that it is difficult to capture the subtle differences between fine-grained categories. In addition, the rough category text descriptions also lead to unsatisfactory performance because they tend to produce text features with lower discriminability. However, as the number of categories increases, the performance degradation of CLIP on any dataset is small. The anti-forgetting effect of most classical FSCIL methods is lower than that of the CLIP method. This powerful anti-forgetting ability is attributed to its extraordinary generalization ability.

[0117] An ablation study was also conducted in this embodiment, and the results on Aircraft100 are reported in Table 3. The following components listed separately in Table 3: I represents the image integration encoder including the CLIP image encoder and the integration of the fine-grained expert model, M represents the feature fusion of image features and text features, called multimodal feature fusion, SC represents the adoption of supervised contrast learning, and F represents the generation of virtual category data based on old category data, called the fantasy method.

[0118] Compared with the original CLIP, the integrated fine-grained expert model (I) has made a breakthrough improvement in AA, with a 28.71% performance increase. This is mainly attributed to the fine-grained expert model, which can effectively learn the differences between fine-grained categories. It enhances the ability of the integrated image encoder to extract more discriminative features. Multimodal feature fusion (M) improves the overall performance of the model. This means that fusing image features with text features generated from detailed text descriptions can significantly enhance the discriminability of the feature prototypes of each category. Supervised contrastive learning (SC) helps to improve the classification performance of old categories to a certain extent. Supervised comparison learning may cause the model to overly prioritize the classification of old categories while ignoring the prediction of new categories that will appear in the future. Therefore, the imagination method (F) becomes crucial, which enables the model to imagine new categories, thereby enhancing the generalization ability. This greatly improves the overall performance of the model.

[0119] Table 3

[0120]

[0121] The method provided in this application utilizes the powerful generalization ability of CLIP to solve the catastrophic forgetting problem and addresses the challenges of FG-FSCIL by integrating a fine-grained expert model and fusing multimodal features. It finally achieved accuracies of 64.93%, 81.09%, 87.72%, and 73.95% on Aircraft100, CUB200, miniImageNet, and CIFAR100, respectively, which are 11.00%, 18.59%, 30.61%, and 22.34% higher than the current SOTA method SAVC. Compared with CLIP, the method provided in this application shows superior performance on the CUB200 and Aircraft100 datasets, with AA values 18.72% and 41.51% higher than CLIP, respectively. It also has some improvements on the two coarse-grained datasets, and on the four datasets, the method provided in this application is comprehensively better than the SOTA model.

[0122] As can be seen from the above, the few-shot class incremental learning method proposed in this application combines the powerful generalization ability of CLIP to solve the catastrophic forgetting problem, uses a fine-grained expert model to assist the CLIP image encoder, and fuses image features and text features to improve the performance of this method in the FG-FSCIL scenario.

[0123] As Figure 7 shown, this application also provides a few-shot class incremental learning device 700, including:

[0124] The first processing module 701 is configured to input the image to be processed into an image integration encoder for encoding to obtain target image features. The image integration encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained categories;

[0125] The second processing module 702 is configured to input the text to be processed into a CLIP text encoder for encoding to obtain target text features;

[0126] The first determination module 703 is configured to determine a learning result based on the target text features and the target image features.

[0127] Optionally, the first processing module 701 includes:

[0128] An input unit, configured to input the image to be processed into the CLIP image encoder to obtain first image features, and input the image to be processed into the fine-grained expert model to obtain second image features;

[0129] A fusion unit, configured to perform fusion processing on the first image features and the second image features to obtain the target image features.

[0130] Optionally, the fusion unit is specifically configured to:

[0131] Input the first image features into a first fully connected layer for dimension adjustment to obtain third image features, and the dimension of the third image features is the same as the dimension of the second image features;

[0132] Average the third image features and the second image features to obtain the target image features.

[0133] Optionally, the few-shot class incremental learning device 700 further includes:

[0134] An iterative training module, configured to iteratively train a to-be-trained integrated network based on a training data set to obtain an integrated network, where the integrated network includes the image integration encoder and the CLIP text encoder;

[0135] Wherein, the training data set includes old category data and virtual category data, and the virtual category data is obtained by performing an augmentation operation on the old category data.

[0136] Optionally, the few-shot class incremental learning device 700 further includes:

[0137] A second determination module, configured to determine a feature prototype of the target category based on the target image features corresponding to the target category and the text features corresponding to the target category.

[0138] Optionally, the few-shot class incremental learning device 700 further includes:

[0139] An acquisition module, configured to acquire supplementary text information corresponding to the target category, where the supplementary text information is used to describe the target category;

[0140] A third processing module, configured to input the supplementary text information into the CLIP text encoder for encoding processing to obtain text features corresponding to the target category.

[0141] Optionally, the fine-grained expert model is a ResNet18 with a self-attention mechanism.

[0142] The few-shot class incremental learning device 700 provided by the embodiments of the present application can execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0143] It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0144] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0145] As Figure 8 shown, an embodiment of the present application provides an electronic device 800, including: a memory 802, a processor 801, and a program stored on the memory 802 and executable on the processor 801; the processor 801 is configured to read the program in the memory 802 to implement the steps in the few-shot class incremental learning method as described above.

[0146] The embodiment of the present application also provides a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements each process of the above-mentioned small-sample class incremental learning method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as compact disks (CD), digital versatile disks (DVD), Blu-ray discs (BD), high-definition versatile discs (HVD), etc.), and semiconductor memories (such as read-only memories (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), non-volatile memories (NAND FLASH), solid-state disks (SSD)).

[0147] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0149] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A small-sample class incremental learning method, characterized in that Including: Input the image to be processed into an image integration encoder for encoding to obtain target image features. The image integration encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained categories; Input the text to be processed into a CLIP text encoder for encoding to obtain target text features; Determine the learning result based on the target text features and the target image features.

2. The method according to claim 1, characterized in that, The step of inputting the image to be processed into an image integration encoder for encoding to obtain target image features includes: Input the image to be processed into the CLIP image encoder to obtain a first image feature, and input the image to be processed into the fine-grained expert model to obtain a second image feature; Fuse the first image feature and the second image feature to obtain the target image features.

3. The method according to claim 2, characterized in that, The step of fusing the first image feature and the second image feature to obtain the target image features includes: Input the first image feature into a first fully connected layer for dimension adjustment to obtain a third image feature, and the dimension of the third image feature is the same as that of the second image feature; Average the third image feature and the second image feature to obtain the target image features.

4. The method according to claim 1, wherein The method further includes: Iteratively train the to-be-trained integrated network based on a training dataset to obtain an integrated network, and the integrated network includes the image integration encoder and the CLIP text encoder; Wherein, the training dataset includes old category data and virtual category data, and the virtual category data is obtained by augmenting the old category data.

5. The method according to claim 4, wherein Before the step of iteratively training the to-be-trained integrated network based on a training dataset to obtain an integrated network, the method further includes: Determine the feature prototype of the target category based on the target image features corresponding to the target category and the text features corresponding to the target category.

6. The method according to claim 5, wherein Before the step of obtaining the feature prototype of the target category based on the target image features corresponding to the target category and the text features corresponding to the target category, the method further includes: Obtain supplementary text information corresponding to the target category, and the supplementary text information is used to describe the target category; Input the supplementary text information into the CLIP text encoder for encoding to obtain the text features corresponding to the target category.

7. The method according to claim 1, characterized in that, The fine-grained expert model is a ResNet18 with a self-attention mechanism.

8. A small-sample class incremental learning device, characterized in that, Including: A first processing module, configured to input the image to be processed into an image integration encoder for encoding to obtain target image features. The image integration encoder includes a CLIP image encoder and a fine-grained expert model, and the fine-grained expert model is used to identify the differences between fine-grained categories; A second processing module, configured to input the text to be processed into a CLIP text encoder for encoding to obtain target text features; A first determination module, configured to determine the learning result based on the target text features and the target image features.

9. An electronic device, comprising: A memory, a processor, and a program stored on the memory and executable on the processor; characterized in that the processor is configured to read the program in the memory to implement the steps in the few-shot class incremental learning method according to any one of claims 1 to 7.

10. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, the steps in the few-shot class incremental learning method according to any one of claims 1 to 7 are implemented.