Small sample class incremental learning method and device

By generating virtual prototypes in incremental learning of small sample classes and separating feature extractors and classifiers, combined with uncertain quantitative selection model, the problem of catastrophic forgetting in incremental learning of small sample classes is solved, achieving higher learning accuracy and effective recognition of new and old categories.

CN120279296APending Publication Date: 2025-07-08HUNAN FIRST NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410286200.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing small sample incremental learning methods are prone to catastrophic forgetting when learning new categories, unable to effectively retain the recognition ability of old categories, and existing methods fail to effectively simulate the memory storage method of the human cerebral cortex.

Method used

Base class training is performed using image data based on base class sessions, virtual prototypes are generated and feature extractors and classifiers are separated, and appropriate classification models are selected through uncertainty quantification to simulate the memory storage mechanism of the human cerebral cortex.

Benefits of technology

It improves the accuracy of incremental learning in small sample classes, effectively retains the ability to identify old categories, reduces catastrophic forgetting when learning new categories, and enhances the model's adaptability to new categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279296A_ABST
    Figure CN120279296A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a small sample class incremental learning method and device, and the method comprises the steps: carrying out the base class training based on image data of a base class session, obtaining a base class model, and enabling the base class model to comprise a feature extractor, a feature extractor tail part, and a classifier; fixing a feature extractor in the base class model, training the tail of the feature extractor and a classifier of the base class model for the small sample image data of the at least one incremental session, and obtaining classification models corresponding to the incremental sessions; inputting an image sample to be classified into the base class model and the classification model of each incremental session, obtaining classification results of the base class model and each classification model, and calculating an uncertainty value of each model; and selecting a classification result of the base class model or the classification model with the minimum uncertainty value as a classification result of the to-be-classified image sample. Through the above mode, the embodiment of the invention provides a new perspective for a small sample class incremental learning domain, and the accuracy of small sample class incremental learning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of few-shot classification, and particularly to a few-shot class incremental learning method and device. Background Art

[0002] Deep learning has achieved milestone achievements in many large-scale computer vision tasks. These methods usually use a large amount of data sets to learn the mapping from samples to corresponding labels. However, the mapping learned from a specific data distribution is usually fixed and non-expandable. That is, the model can only recognize the trained categories and cannot be generalized to new categories. In order to enable the trained model to be effectively generalized to new categories, class incremental learning has received extensive attention. Class incremental learning aims to enable the model to continuously learn different categories from a data stream rather than from a fixed data set, while retaining the recognition ability for previously encountered categories. Most class incremental learning studies focus on enabling the model to continuously learn when new category samples are sufficient. However, in many practical scenarios, obtaining samples from new categories may be challenging and costly. To address the challenge of scarce samples from new categories, a more challenging and practical task has been proposed, called few-shot class incremental learning. Compared with class incremental learning, few-shot class incremental learning aims to perform incremental learning in the case where the labeled samples of new categories are extremely limited. Current few-shot class incremental learning studies mainly follow the traditional class incremental learning paradigm, that is, a single model learns all data during the entire incremental process, sharing the same model parameters and decision boundaries. This learning paradigm faces significant challenges in retaining the old memories of the model. When acquiring knowledge of new categories, it inevitably changes the parameters trained on old categories, resulting in catastrophic forgetting, deviating from the way the human brain stores memories. Summary of the Invention

[0003] In view of the above problems, the embodiments of the present invention provide a few-shot class incremental learning method and device, which overcome or at least partially solve the above problems.

[0004] According to one aspect of an embodiment of the present invention, a small-sample class incremental learning method is provided. The method includes: performing base class training on the image data of the base class session to obtain a base class model, where the base class model includes a feature extractor, a tail of the feature extractor, and a classifier; fixing the feature extractor in the base class model, and training the tail of the feature extractor and the classifier of the base class model with the small-sample image data of at least one incremental session to obtain classification models corresponding to each incremental session respectively; inputting the image sample to be classified into the base class model and the classification models of each incremental session, obtaining the classification results of the base class model and each classification model, and calculating the uncertainty value of each model; and selecting the classification result of the base class model or the classification model with the smallest uncertainty value as the classification result of the image sample to be classified.

[0005] Optionally, the performing base class training on the image data of the base class session to obtain a base class model includes: using the CutMix data augmentation method to generate virtual samples for the real class samples in the image data of the base class session; extracting feature vectors from the generated virtual samples and synthesizing virtual prototypes; performing preliminary training on the base class model based on the image data of the base class session and the virtual prototypes, and after the preliminary training is completed, adjusting the number of neurons in the tail of the fully connected layer of the model to match the number of real classes; and fine-tuning the trained base class model with the samples of real classes.

[0006] Optionally, the using the CutMix data augmentation method to generate virtual samples for the real class samples in the image data of the base class session includes: for samples with similar sample feature vectors, using the CutMix data augmentation method to crop and mix the image X A onto the image X B to obtain a virtual class image obtained by fusing the image X A and the image X B as a virtual sample:

[0007]

[0008] where is the virtual class image, M is a binary mask for masking image information, ⊙ represents the pixel-by-pixel multiplication of two images, and H and W are the height and width of the image respectively.

[0009] Optionally, to fix the feature extractor in the base class model and train the tail of the feature extractor and the classifier of the base class model with small sample image data of at least one incremental session to obtain a classification model corresponding to each incremental session, including: applying the fixed feature extractor to extract features from each training sample of any category k in the small sample image data of at least one incremental session to generate a feature vector P i ; calculating the mean of the feature vectors P i extracted from each training sample of any category k to obtain a feature prototype Fine-tune the tail of the feature extractor and the classifier of the base class model according to the cross-entropy loss function to obtain a classification model corresponding to each incremental session.

[0010] Optionally, after inputting the image sample to be classified into the base class model and each of the classification models, obtaining the classification results of the base class model and the classification models of each incremental session, and calculating the uncertainty value of each model, it further includes: inputting the image sample to be classified into the base class model and the classification models {M 0 , M 1 , …, M n} of each incremental session, outputting classification results {R 0 (x), R 1 (x), …, R n (x)}, and using information entropy to measure the uncertainty of the model with respect to the image sample to be classified:

[0011]

[0012] where H t (x) represents the uncertainty value of the classification model of the t-th incremental session with respect to the image sample X to be classified, |C t | represents the number of categories in the t-th incremental session, and p(o c ) represents the probability that the image sample X to be classified is of category c.

[0013] Optionally, after inputting the image sample to be classified into the base class model and each of the classification models, obtaining the classification results of the base class model and the classification models of each incremental session, and calculating the uncertainty value of each model, it includes: dividing the classification result of the base class model into multiple sub-class results according to the target category of the classification model of the incremental session, and obtaining the uncertainty value corresponding to each sub-class result; selecting the classification result corresponding to the minimum uncertainty value from the uncertainty values corresponding to each sub-class result as the classification result of the base class model for the image sample to be classified.

[0014] Optionally, dividing the classification result of the base class model into multiple subclass results according to the target category of the classification model of the incremental session, including: applying the following relational expression according to the target category of the classification model of the incremental session to divide the classification result of the base class model into N sub subclass results:

[0015]

[0016] wherein, |C 0 | and |C i | respectively represent the number of categories of the base class session and the incremental session.

[0017] Based on the same inventive concept, a few-shot class incremental learning device is provided, including: a base class model acquisition unit, configured to perform base class training based on the image data of the base class session to obtain a base class model, where the base class model includes a feature extractor, a tail of the feature extractor, and a classifier; a classification model acquisition unit, configured to fix the feature extractor in the base class model, and train the tail of the feature extractor and the classifier of the base class model with the small sample image data of at least one incremental session to obtain classification models respectively corresponding to the incremental sessions; a classification prediction evaluation unit, configured to input an image sample to be classified into the base class model and the classification models of the incremental sessions, obtain the classification results of the base class model and the classification models and calculate the uncertainty value of each model; a classification result determination unit, configured to select the classification result of the base class model or the classification model with the smallest uncertainty value as the classification result of the image sample to be classified.

[0018] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the foregoing method when executing the program.

[0019] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium, in which at least one executable instruction is stored, and the executable instruction enables the processor to execute the foregoing method.

[0020] In an embodiment of the present invention, base training is performed on image data based on a base class session to obtain a base class model, where the base class model includes a feature extractor, a tail of the feature extractor, and a classifier; the feature extractor in the base class model is fixed, and the tail of the feature extractor and the classifier of the base class model are trained with small sample image data of at least one incremental session to obtain classification models respectively corresponding to the incremental sessions; the image sample to be classified is input into the base class model and the classification models of the incremental sessions, the classification results of the base class model and the classification models are obtained, and the uncertainty value of each model is calculated; the classification result of the base class model or the classification model with the smallest uncertainty value is selected as the classification result of the image sample to be classified, providing a new perspective for the small sample class incremental learning domain and being able to improve the accuracy of small sample class incremental learning.

[0021] The above description is only an overview of the technical solution of the embodiment of the present invention. In order to be able to understand the technical means of the embodiment of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the embodiment of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0023] Figure 1 A flowchart showing the small sample class incremental learning method provided by the embodiment of the present invention is shown;

[0024] Figure 2 A schematic diagram of the overall framework of the small sample class incremental learning method provided by the embodiment of the present invention is shown;

[0025] Figure 3 A schematic diagram of the virtual prototype application of the small sample class incremental learning method provided by the embodiment of the present invention is shown;

[0026] Figure 4 A schematic diagram of the branch training of the small sample class incremental learning method provided by the embodiment of the present invention is shown;

[0027] Figure 5 A schematic diagram of the performance comparison of the small sample class incremental learning method provided by the embodiment of the present invention on three benchmark datasets is shown;

[0028] Figure 6 A schematic diagram of the structure of the small sample class incremental learning device provided by the embodiment of the present invention is shown;

[0029] Figure 7 The schematic diagram of the electronic device in the embodiment of the present invention is shown. Specific embodiments

[0030] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0031] Compared with the learning paradigm with heavy memory burden and easy to forget, the human brain stores the learned knowledge in different regions of the cerebral cortex. After encountering a certain task, humans can quickly associate it with the corresponding region of the cerebral cortex. This adaptive learning method is particularly suitable for few-shot class incremental learning because it can store memories separately. However, from the perspectives of biology and cognitive science, the challenges of deep neural networks in mimicking the human cerebral cortex in few-shot class incremental learning are mainly reflected in the following two aspects. i) How can the model imitate the cerebral cortex to store the acquired knowledge in a partitioned manner? Since the session information in the test phase is unknown, the model must have the ability to recognize all previously learned categories. This limitation results in the model having to store all the knowledge obtained in all sessions in an inseparable unified memory region. ii) How to establish the mapping relationship from tasks to memories? In few-shot class incremental learning, it is unknown which session the test samples come from, so it becomes infeasible to establish the mapping relationship between the samples and the model. This means that directly mapping the relevant regions of the cerebral cortex, just as the human brain does naturally for a given task, is infeasible for deep neural networks. The embodiments of the present invention are committed to bridging the gap between the few-shot class incremental learning process and the memory storage method of the human cerebral cortex by solving the above challenges.

[0032] Figure 1 The schematic flow chart of the few-shot class incremental learning method provided by the embodiment of the present invention is shown. As Figure 1 shown, this few-shot class incremental learning method is applied to a server, and the few-shot class incremental learning method includes:

[0033] Step S11: Perform base class training based on the image data of the base class session to obtain a base class model, where the base class model includes a feature extractor, a feature extractor tail, and a classifier.

[0034] In the embodiments of the present invention, the deep neural network ResNet is used as the backbone network for classification in few-shot class incremental learning. For example, ResNet-18 is selected as the backbone network. Few-shot class incremental learning is a more challenging learning task compared to class incremental learning. Few-shot class incremental learning aims to continuously learn new classes with limited labeled data while preventing forgetting the previously learned classes. The problem setting of few-shot class incremental learning was first proposed in TOPIC, which uses a neural gas network to stabilize the topological manifold features between new and old classes. CEC uses a graph neural network to update the classifier parameters in each incremental session based on the global knowledge of the previous session. At the same time, during the base class training process, it creates pseudo-incremental scenarios to optimize the graph neural network. Different from the previous methods that only consider the performance of the current session, it enhances the incremental model scalability by compressing the embedding space through virtual prototypes. In the embodiments of the present invention, images with different classes are cropped and mixed to form images of virtual classes, and virtual prototypes are generated relying on semantic information similar to real classes.

[0035] In the embodiments of the present invention, the fundamental goal of few-shot class incremental learning is to continuously learn new classes in a small number of labeled samples while ensuring that the previously acquired knowledge is retained. In few-shot class incremental learning, the complete training dataset D train is represented as where represents the training data of the t-th session, and N represents the number of sessions. The associated label space is The samples and labels of different sessions do not overlap. is the base training dataset, which contains a large number of classes and each class has sufficient training samples. is the training dataset of the incremental session, which only contains a small number of samples. Specifically, it contains N classes, and each class contains K samples. During the training of the t-th incremental session, the model can only access the data and labels of the current session, and In the test phase of the t-th incremental session, evaluate the classes that the model has learned

[0036] For the base-class session training, a data augmentation method based on CutMix is used to generate virtual prototypes for training a more robust feature extractor. For incremental sessions, an independent classification model is trained for each incremental session, and the classification model for each incremental session only retains the memory of the current session to avoid forgetting. This is the first parameter isolation method in Few-Shot Class-Incremental Learning (FSCIL). To prevent overfitting to limited training data during incremental learning, a strategy of fine-tuning the tail structure of the model instead of the entire model is adopted. In the testing phase, to help the sample select an appropriate classification model, the uncertainty of the input sample is quantified by calculating the information entropy of the sample, enabling the sample to select the most suitable classification model. As Figure 2 shown, applying the CutMix data augmentation method enhances the performance of the feature extractor by generating virtual prototypes. For the t-th session, the main part F of the feature extractor is frozen Φ , and the tail and the classifier W are fine-tuned t . In the testing phase, the sample X is input into a series of trained models, and the classification result R t (X) and the uncertainty H t (X) are output, and then a reliable result is selected based on the value of H t (X).

[0037] Few-Shot Class-Incremental Learning usually uses a metric learning method based on the nearest neighbor algorithm to classify samples, and the class of the prototype closest to the sample feature vector is predicted as the class to which the sample belongs. However, as Figure 3 shown in a, if the prototypes in the embedding space are too close, the error rate of sample classification will increase. To increase the interval between prototypes and create a sparser embedding space, a method of generating virtual prototypes is used. This method inserts the generated virtual prototypes into the embedding space. In step S11, optionally, first the CutMix data augmentation method is used to generate virtual samples for the real-class samples in the image data based on the base-class session; then feature vectors are extracted from the generated virtual samples, and virtual prototypes are synthesized. Specifically, as Figure 3 shown in c, for samples with close sample feature vectors, the CutMix data augmentation method is used to crop and mix the image X A onto the image X B to obtain the virtual-class image obtained by fusing the image X A and the image X B as the virtual sample:

[0038]

[0039] Among them, is the virtual class image obtained by fusing image X A and image X B The virtual class image is obtained by fusing image X. M is a binary mask used to mask image information. ⊙ represents the per-pixel multiplication of two images. H and W are the height and width of the image respectively. This method is called CutMix. Then, feature vectors are extracted from these virtual samples to synthesize virtual prototypes. In the embodiments of the present invention, the size of M is restricted to half of the original image to ensure the comprehensive integration of the semantic features of the two classes into the virtual class. As Figure 3 shown in Fig. b, the virtual samples generated by the CutMix method do not have the existence of non-pixel information, enabling the virtual prototypes to be closely aligned with the real classes in the embedding space, further enhancing the interval between the real prototypes, and enabling the feature extractor to extract more discriminative class features.

[0040] After obtaining the virtual prototypes, the base class model is trained based on the image data of the base class session and the virtual prototypes, and the number of neurons at the tail of the fully connected layer of the model is adjusted to match the number of real classes; the trained base class model is fine-tuned using samples of real classes. The generation of virtual samples enhances the feature extraction ability of the backbone. However, it simultaneously introduces non-existent classes into the classification system. These non-existent classes may disrupt the stability of the classification results. This effect may make the results of uncertainty quantification unreliable. To enhance the reliability of the uncertainty quantification results, the model can be fine-tuned using real classes. After training a powerful backbone using samples containing virtual classes, the number of neurons at the tail of the fully connected layer is adjusted to match the number of real classes in the current session. Subsequently, the model is fine-tuned using training samples containing only real classes. For the base session, the entire network is fine-tuned because this session contains a large number of samples. In contrast, in the incremental session, due to the limited number of samples, the fine-tuning only involves the classifier.

[0041] Step S12: Fix the feature extractor in the base class model, and train the tail of the feature extractor and the classifier of the base class model with the small-sample image data of at least one incremental session to obtain classification models corresponding to each incremental session respectively.

[0042] Few-shot learning aims to identify unseen objects from a small number of labeled samples. Currently, it can be mainly divided into metric-based methods, optimization-based methods, and model-based methods. Metric-based methods use the nearest neighbor algorithm to learn the relationship between sample pairs and assign labels to the samples with the smallest distance by calculating the distance between sample pairs. For example, the mean of all samples in each class is used to represent the position of each class in the embedding space. In the embodiments of the present invention, features are first extracted using a pre-trained backbone network, and then classification is performed by calculating the cosine similarity between the samples and the prototypes in the embedding space.

[0043] In the field of class-incremental learning, the model continuously learns new knowledge from a data stream with multiple classes. Embodiments of the present invention focus on class-incremental learning in the case of limited new class samples. For example, iCaRL uses a nearest-neighbor classifier to learn new classes and adopts knowledge distillation as a strategy to retain the knowledge of existing classes. LUCIR improves performance in subsequent sessions by searching for flat minima during the base training process. The following assumptions can be made about LUCIR: the base-class model plays a key role in the overall performance of the incremental task. It builds a powerful feature extractor by training on the base classes and then freezes it. Then, for each incremental session, it fine-tunes the tail layers of the feature extractor and the classifier separately. Inspired by this, embodiments of the present invention adopt the strategy of fine-tuning the tail of the feature extractor in few-shot class-incremental learning (FSCIL) for incremental learning. Compared with completely freezing the feature extractor, this strategy enables the model to better adapt to new classes.

[0044] To suppress catastrophic forgetting in the incremental process of deep neural networks, the strategy of freezing the parameters of the feature extractor after base-class training is adopted. This method effectively suppresses forgetting but imposes a limitation on the ability of the feature extractor to learn new classes. To facilitate the adjustment of model parameters on new classes while retaining the feature extraction ability obtained on the base classes, the feature extractor is divided into two parts, F Φ and F e 。F Φ is the main component of the network, F e is the smaller layer at the end of the network. In the t-th incremental session, the entire classification model M includes F Φ , and W t ,W t is the classifier for the t-th incremental session. The feature vector P i extracted from the image X i by the feature extractor. The feature extractor extracts the feature vectors of all training samples of class k and calculates the mean to obtain the feature prototype of this class

[0045] In embodiments of the present invention, the fixed feature extractor is applied to extract features from each training sample of any class k in the few-shot image data of at least one incremental session, generating a feature vector P i :

[0046] P i =F e (F φ (x i ))

[0047] Calculate the mean of the feature vectors P extracted from each training sample of any category k to obtain the feature prototype i where |C(k)| represents the total number of training samples in category k,

[0048]

[0049] is the feature vector of the i-th sample in category k. Then, according to the cross-entropy loss function, fine-tune the tail of the feature extractor and the classifier of the base class model to obtain classification models corresponding to each incremental session.

[0050] The difference between few-shot class incremental learning and traditional class incremental learning is that the new categories in few-shot class incremental learning only contain extremely limited training data. Therefore, adjusting Fin incremental training Φ may lead to severe overfitting to the new categories. Therefore, the parameters of F Φ are shared among all session models. The specific training process is as Figure 4 shown. When there is enough data in the base class session, train F Φ , and W 0 . In the incremental session, only fine-tune and W t , and the parameters of F Φ are frozen. This training strategy suppresses the forgetting phenomenon and enables the model to better adapt to new knowledge. In the base class session and the incremental session, use the cross-entropy loss function shown by the following relational expression for training:

[0051]

[0052] where L CE is the value of the cross-entropy loss function, is the feature vector of the j-th sample in category k, and N represents the total number of samples in category k.

[0053] The methods for obtaining the corresponding classification models based on the incremental session include but are not limited to iCaRL, EEIL, LUCIR, TOPIC, CEC, F2M, Entropy-reg, MetaFSCIL, GKEAL.

[0054] Step S13: Input the image samples to be classified into the base class model and the classification models of each incremental session, obtain the classification results of the base class model and each classification model, and calculate the uncertainty value of each model.

[0055] Since session information is not accessible during the testing of few-shot class-incremental learning, establishing a mapping from samples to the correct model becomes a key challenge. During incremental learning, the tail of the feature extractor and the classifier for each session are trained independently. Thus, there is a feature extractor backbone F Φ , a series of feature extractor tails and a series of classifiers {W 0 , W 1 , …, W n}. It is crucial to ensure that samples are input into the corresponding session models to guarantee the validity of the output classification results. Uncertainty quantification techniques can be used to measure how familiar the model is with the input samples. The classification model generates an uncertainty value through uncertainty quantification while outputting the classification result. Low uncertainty indicates that the model recognizes the sample; conversely, high uncertainty indicates that the model lacks confidence in successfully identifying this sample. Information entropy can be used to measure the uncertainty of the model with respect to the input samples.

[0056] In an embodiment of the present invention, the image sample to be classified is input into the base class model and each of the classification models {M 0 , M 1 , …, M n}, and the classification results {R 0 (x), R 1 (x), …, R n (x)} are output, and information entropy is used to measure the uncertainty of the model with respect to the image sample to be classified:

[0057]

[0058] where H t (x) represents the uncertainty value of the classification model of the t-th incremental session with respect to the image sample X to be classified, |C t | represents the number of classes in the t-th incremental session, and p(o c ) represents the probability that the image sample X to be classified is of class c.

[0059] To measure the uncertainty of samples in different classification systems, it is crucial to ensure that the target classes of different classification systems are the same. This standardization makes the results of uncertainty more meaningful and comparable for analysis. For example, in CIFAR-100, the model of the base session selects one class from 60 classes, while the model of the incremental session selects one class from 5 classes. The degrees of uncertainty of these two models with respect to the samples are obviously different. To address the class imbalance problem in model selection based on uncertainty, the uncertainty quantification results of the base session model are further refined. To make the target class of the base class model M 0 the same as that of the incremental session model M tThe target category is consistent, and M can be based on the number of target categories of t M to divide the output result of 0 .

[0060] In the embodiment of the present invention, according to the target category of the classification model of the incremental session, the classification result of the base model is divided into multiple sub-results, and the uncertainty value corresponding to each sub-result is obtained. Specifically, according to the target category of the classification model of the incremental session, the following relational expression is applied to divide the classification result of the base model into N sub sub-results:

[0061]

[0062] wherein, |C 0 | and |C i | respectively represent the number of categories of the base session and the incremental session. In this way, the classification result of the base model M 0 is divided into N sub sub-results, and each sub-result includes the output values of the classifier for |C i | categories.

[0063] Then, select the classification result corresponding to the smallest uncertainty value from the uncertainty values corresponding to each sub-result as the classification result of the base model for the image sample to be classified. Use the uncertainty value of the sub-result to participate in model selection, rather than relying on the uncertainty value of the overall classification result of the base model M 0 . Specifically, in the test phase, each sample X i is input into the models {M 0 , M 1 , …, M n} of all sessions, and a set of uncertainty values shown in the following formula is generated:

[0064] U = {H 0 (x i ), H 1 (x i ), …, H n (x i )}

[0065] wherein, U is the uncertainty set, and H 0 (x i ) represents the uncertainty value of the base model for the image sample x i to be classified, and H j (x i ) represents the uncertainty value of the classification model of the incremental session j for the image sample x i to be classified.

[0066] To ensure the unity of the target category, M 0 's N sub subclass results are incorporated into the uncertainty set, rather than H 0 (x i ). This is because the output result of M 0 |C 0 | contains more categories than |C i |. Therefore, incorporating the subclass results of M 0 into the uncertainty quantification ensures that the number of categories is consistent with |C i |. The uncertainty set obtained after inputting the samples into all models is shown by the following formula:

[0067]

[0068] where, represents the set of all sub-results from the base class model M 0 . This strategy of splitting the sub-results of the base session model effectively solves the problem of class imbalance between the base class and the incremental class, achieving this goal by aligning the number of categories in the subclass results of the base class model with the incremental class.

[0069] Step S14: Select the classification result of the base class model or the classification model with the smallest uncertainty value as the classification result of the to-be-classified image sample.

[0070] In the embodiment of the present invention, the model with the smallest uncertainty is selected as the classification model of the to-be-classified image sample, and the classification result of this classification model is considered as the final classification result of the to-be-classified image sample. If the smallest uncertainty value is located in , then the base class model is selected as the classification model of the to-be-classified image sample, and the input sample X i selects the classification result corresponding to this minimum value as the predicted category of the base class model M 0 for the input sample X i .

[0071] The performance of the few-shot class incremental learning method according to the embodiments of the present invention is illustrated by the following examples. The datasets applied include CIFAR-100, Mini-ImageNet, and CUB-200. CIFAR-100 contains 100 classes, with 600 32×32 RGB images for each class, where 500 are for training and 100 are for testing, and the entire dataset has 600,000 images. Among them, 60 classes are base classes and 40 classes are incremental classes, and there are eight sessions in the incremental stage. The data format in each incremental session is 5-way 5-shot, indicating a few-shot image classification task of 5-classification, and there are 5 image samples for each class to train the model. Mini-ImageNet is a subset of ImageNet, containing a total of 100 classes, with 600 84×84 RGB images for each class. Among them, 60 classes are used as base classes and 40 classes are used as incremental classes. These 40 classes are evenly distributed into eight sessions, with five classes in each session. The data in each incremental session is presented as 5-way 5-shot. CUB-200 contains 11,788 224×224 RGB images from 200 classes. Among them, 100 classes are used as base classes and another 100 classes are used as incremental classes. The incremental classes are divided into 10 sessions, with 10 classes in each session. The data in each incremental session is presented as 10-way 5-shot. ResNet-18 is selected as the backbone network to verify the performance of the few-shot class incremental learning method according to the embodiments of the present invention. For CIFAR-100 and ImageNet, randomly initialized model parameters are used. For CUB-200, pre-trained model parameters are used. The feature extraction part of the entire ResNet-18 is divided into four blocks, and each block contains three convolutional layers. In our branch training strategy, the first three blocks are F Φ , and the parameters are frozen after the base class session training. The last block at the tail is F e , and the parameters are continuously updated during the incremental session. The following recent methods are used as baselines in the comparative experiment: iCaRL, EEIL, LUCIR, TOPIC, CEC, F2M, Entropy-reg, MetaFSCIL, GKEAL. The performance reports of these baselines are from GKEAL for fair comparison. In the ablation experiment, if there is no branch training, it means that only the classifier is fine-tuned during the incremental process, while the parameters of the feature extractor remain frozen.

[0072] The optimizer uses Stochastic Gradient Descent (SGD) with a learning rate of 0.1 for the base class session and SGD with a learning rate of 0.05 for the incremental session. The momentum in SGD is set to 0.9. For CIFAR-100, the number of training epochs is 200 in the base class session and 20 in the incremental session. For mini-ImageNet, the number of training epochs is 300 in the base class session and 30 in the incremental session. For CUB-200, the number of training epochs is 180 in the base class session and 20 in the incremental session.

[0073] To verify the state-of-the-art performance of the method of the embodiment of the present invention, comparative experiments with existing methods were conducted on three benchmark datasets: CIFAR-100, mini-ImageNet, and CUB-200. The comparative performance of the method of the embodiment of the present invention with other FACIL methods on the three benchmark datasets is as Figure 5 shown. In addition, the detailed performance of each session of CIFAR-100 is shown in Table 1. Here, 0 represents the base class session, 2-8 represent the incremental sessions, and AA represents the average value.

[0074] Table 1 Comparison of the accuracies of each method on each session on CIFAR-100

[0075]

[0076]

[0077] According to the experimental results, the method of the embodiment of the present invention has achieved satisfactory performance on both CIFAR-100 and mini-ImageNet, not only performing well in the base session but also in each subsequent incremental session. In CIFAR-100, the method of the embodiment of the present invention is 3.33% higher than GKEA in the last session, and the average accuracy has increased by 6.09%. The experimental results clearly show that the method of the embodiment of the present invention outperforms other state-of-the-art FSCIL methods on natural image datasets. The parameter-independent method proposed in FSCIL effectively alleviates the catastrophic forgetting problem. This is achieved by simulating the human cerebral cortex and using different models to store the knowledge obtained from different sessions. At the same time, the experimental results prove that the uncertainty quantification technology can construct a mapping from samples to models. This learning method that simulates the human cerebral cortex provides a new perspective for the entire FSCIL field. However, the method of the embodiment of the present invention performs poorly on CUB-200, only performing better than iCaRL, EEIL, LUCIR, and TOPIC. This deficiency will be analyzed in detail in the following section.

[0078] To evaluate the effectiveness of the method, ablation experiments were conducted on CIFAR-100. The detailed experimental results are shown in Table 2, where AA represents the average accuracy, CM represents the virtual prototype generation method, BR represents the branch training strategy, MS represents the model selection method, SR represents the sub-result of the base class session model, and FT represents the fine-tuning according to the actual number of classes. After applying the virtual prototype generation method, the accuracy was improved to 82.93%. This indicates that the virtual prototype generation method is effective in enhancing the feature extraction ability of the backbone network. Based on the enhancement of the base class session performance, the branch training strategy was implemented to train models for each session. Subsequently, the model selection method based on uncertainty quantification was used to determine the most suitable classification result for a given sample during the test phase. At the same time, the model of the base class session uses the sub-result of uncertainty quantification to solve the class imbalance problem. The experimental results show that in the incremental session phase, the method of simulating the human cerebral cortex in the embodiments of the present invention produced a higher average accuracy compared to the traditional non-parametric independent method, with an increase of 1.37%. Under the same base session performance, the effectiveness of the final session was improved, from 49.37% to 51.79%. Fine-tuning the classifier according to the actual number of classes in each session can significantly improve the accuracy of uncertainty quantification. Fine-tuning the classifier increased the average accuracy from 65.27% to 67.44%. It is worth noting that the accuracy of the final session showed a more significant improvement, from 51.40% to 54.73%. Therefore, the virtual classes that do not exist in the classification model will interfere with the uncertainty quantification process of the model.

[0079] Table 2 Ablation Experiments on CIFAR-100

[0080]

[0081] In the embodiments of the present invention, to address the first challenge of enabling the classification system to store memories in regions like the cerebral cortex, a parameter-independent method is adopted to train independent classification models for each session. In traditional few-shot class incremental learning, the knowledge of all sessions is integrated into a single model. This integration poses a risk of catastrophic forgetting. In few-shot class incremental learning based on parameter independence, each model specializes in storing the knowledge obtained in its corresponding session, and the decision boundaries of these models are independent. This method of retaining prior knowledge effectively reflects the memory partition storage mechanism in the human cerebral cortex. Additionally, to mitigate overfitting to new classes, the embodiments of the present invention use a branch training strategy to train a series of models. To address the second challenge of the mapping from samples to models, session information of samples is predicted through uncertainty quantification. Test samples are individually input into each model to obtain classification results and corresponding uncertainty values. Then, appropriate classification results are selected based on the uncertainty values. Uncertainty quantification enables the model to not only generate its classification results but also express the level of uncertainty associated with a given sample. The model gives low uncertainty for the classes it recognizes and high uncertainty for the classes it has never encountered. Relying on uncertainty quantification, a memory mapping similar to the cerebral cortex is successfully constructed. For each test sample, models that retain the memories of the corresponding sessions are screened out from a series of models.

[0082] The few-shot class incremental learning method of the embodiments of the present invention simulates the memory storage mechanism of the human cerebral cortex from a new perspective of few-shot class incremental learning. It is the first parameter-independent method in few-shot class incremental learning and achieves state-of-the-art performance on benchmark datasets. Independent classification models are trained for each session to maintain the performance of the entire classification system for old classes; a novel data augmentation strategy is applied to generate virtual prototypes for real classes, significantly enhancing the feature extraction ability of the backbone network; in the test stage, by quantifying the uncertainty of samples, session information is obtained, facilitating the selection of appropriate classification models. By simulating the memory storage mechanism of the cerebral cortex, the learning process of few-shot class incremental learning is made more in line with human learning habits, providing a new perspective for the field of few-shot class incremental learning.

[0083] In summary, the few-shot class incremental learning method according to the embodiments of the present invention performs base class training on the image data based on the base class sessions to obtain a base class model, where the base class model includes a feature extractor, a tail of the feature extractor, and a classifier; fixes the feature extractor in the base class model, and trains the tail of the feature extractor and the classifier of the base class model with the few-shot image data of at least one incremental session to obtain classification models corresponding to the respective incremental sessions; inputs the image sample to be classified into the base class model and the classification models of the respective incremental sessions, obtains the classification results of the base class model and the respective classification models, and calculates the uncertainty value of each model; selects the classification result of the base class model or the classification model with the smallest uncertainty value as the classification result of the image sample to be classified, providing a new perspective for the few-shot class incremental learning domain and being able to improve the accuracy of few-shot class incremental learning.

[0084] The specific embodiments of the present invention have been described above. In some cases, the actions or steps recorded in the embodiments of the present invention may be executed in an order different from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0085] Based on the same concept, the embodiments of the present invention further provide a few-shot class incremental learning device. It is applied to a server. As Figure 6 shown, the few-shot class incremental learning device includes: a base class model acquisition unit, a classification model acquisition unit, a classification prediction evaluation unit, and a classification result determination unit. Among them,

[0086] The base class model acquisition unit is configured to perform base class training on the image data based on the base class sessions to obtain a base class model, where the base class model includes a feature extractor, a tail of the feature extractor, and a classifier;

[0087] The classification model acquisition unit is configured to fix the feature extractor in the base class model, and train the tail of the feature extractor and the classifier of the base class model with the few-shot image data of at least one incremental session to obtain classification models corresponding to the respective incremental sessions;

[0088] The classification prediction evaluation unit is configured to input the image sample to be classified into the base class model and the classification models of the respective incremental sessions, obtain the classification results of the base class model and the respective classification models, and calculate the uncertainty value of each model;

[0089] The classification result determination unit is configured to select the classification result of the base class model or the classification model with the smallest uncertainty value as the classification result of the image sample to be classified.

[0090] For convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the embodiments of the present invention, the functions of each module can be implemented in one or more software and / or hardware.

[0091] The device in the above embodiment is applied to the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated herein.

[0092] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in any one of the above embodiments.

[0093] An embodiment of the present invention provides a non-volatile computer storage medium, which stores at least one executable instruction, and the computer executable instruction can execute the method described in any one of the above embodiments.

[0094] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 701, a memory 702, an input / output interface 703, a communication interface 704, and a bus 705. Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are communicatively connected to each other inside the device through the bus 705.

[0095] The processor 701 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the method embodiments of the present invention.

[0096] The memory 702 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 702 may store an operating system and other application programs. When implementing the technical solutions provided by the method embodiments of the present invention through software or firmware, the relevant program codes are stored in the memory 702 and called and executed by the processor 701.

[0097] The input / output interface 703 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices can include a display, a speaker, a vibrator, an indicator light, etc.

[0098] The communication interface 704 is used to connect to the communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. The communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.).

[0099] The bus 705 includes a path for transmitting information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704).

[0100] It should be noted that although the above device only shows the processor 701, the memory 702, the input / output interface 703, the communication interface 704, and the bus 705, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present invention, and do not have to include all the components shown in the figure.

[0101] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of brevity.

[0102] This application aims to cover all such substitutions, modifications, and variations that fall within the broad scope of all embodiments. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the present disclosure.

Claims

1. A small-sample class incremental learning method, characterized in that, The method includes: Performing base-class training on the image data based on the base-class session to obtain a base-class model, where the base-class model includes a feature extractor, a tail of the feature extractor, and a classifier; Fixing the feature extractor in the base-class model, and training the tail of the feature extractor and the classifier of the base-class model with small-sample image data of at least one incremental session to obtain classification models respectively corresponding to each incremental session; Inputting the image sample to be classified into the base-class model and the classification models of each incremental session, obtaining the classification results of the base-class model and each classification model, and calculating the uncertainty value of each model; Selecting the classification result of the base-class model or the classification model with the smallest uncertainty value as the classification result of the image sample to be classified.

2. The method according to claim 1, characterized in that, The performing base-class training on the image data based on the base-class session to obtain a base-class model includes: Using the CutMix data augmentation method to generate virtual samples for the real-class samples in the image data based on the base-class session; Extracting feature vectors from the generated virtual samples and synthesizing virtual prototypes; Performing preliminary training on the base-class model based on the image data of the base-class session and the virtual prototypes. After the preliminary training is completed, adjusting the number of neurons in the tail of the fully connected layer of the model to match the number of real classes; Fine-tuning the trained base-class model with samples of real classes.

3. The method according to claim 2, characterized in that, The using the CutMix data augmentation method to generate virtual samples for the real-class samples in the image data based on the base-class session includes: For samples with similar sample feature vectors, use the CutMix data augmentation method to crop and mix the image X A onto the image X B to obtain the image X A and the image X B The fused virtual class image is used as a virtual sample: Among them, is a virtual category image, M is a binary mask used to conceal image information, ⊙ represents the pixel-by-pixel multiplication of two images, and H and W are the height and width of the image, respectively.

4. The method according to claim 1, characterized in that, The fixing the feature extractor in the base-class model, and training the tail of the feature extractor and the classifier of the base-class model with small-sample image data of at least one incremental session to obtain classification models respectively corresponding to each incremental session includes: Use the fixed feature extractor to extract features from each training sample of any category k in the small sample image data of at least one incremental session, generating a feature vector P i ; Calculate the mean of the feature vectors P extracted from each training sample of any category k to obtain a feature prototype i ​ Performing training fine-tuning on the tail of the feature extractor and the classifier of the base-class model according to the cross-entropy loss function to obtain classification models respectively corresponding to each incremental session.

5. The method according to claim 1, wherein The inputting the image sample to be classified into the base-class model and the classification models of each incremental session, obtaining the classification results of the base-class model and each classification model, and calculating the uncertainty value of each model further includes: Input the image sample to be classified into the base class model and the classification models {M 0 ,M 1 ,…,M n} of each incremental session, and output the classification results {R 0 (x),R 1 (x),…,R n (x)}, and use information entropy to measure the uncertainty of the model for the image sample to be classified: Among them, H t (x) represents the uncertainty value of the classification model of the t-th incremental session for the image sample X to be classified, |C t | represents the number of classes in the t-th incremental session, and p(o c ) represents the probability that the image sample X to be classified is of class c.

6. The method according to claim 1, wherein After the inputting the image sample to be classified into the base-class model and the classification models of each incremental session, obtaining the classification results of the base-class model and each classification model, and calculating the uncertainty value of each model, it includes: Dividing the classification result of the base-class model into multiple sub-class results according to the target class of the classification model of the incremental session, and obtaining the uncertainty value corresponding to each sub-class result; Selecting the classification result corresponding to the smallest uncertainty value from the uncertainty values corresponding to each sub-class result as the classification result of the base-class model for the image sample to be classified.

7. The method according to claim 6, wherein The dividing the classification result of the base-class model into multiple sub-class results according to the target class of the classification model of the incremental session includes: Apply the following relational expression according to the target category of the classification model of the incremental session to divide the classification result of the base class model into N sub sub-class results: Among them, |C 0 | and |C i | respectively represent the number of categories of the base class session and the incremental session.

8. A small-sample class incremental learning device, characterized in that The device includes: The base class model acquisition unit is used to perform base class training based on the image data of the base class session and obtain a base class model, where the base class model includes a feature extractor, a tail of the feature extractor, and a classifier; The classification model acquisition unit is used to fix the feature extractor in the base class model and train the tail of the feature extractor and the classifier of the base class model with the small sample image data of at least one incremental session to obtain a classification model corresponding to each incremental session; The classification prediction evaluation unit is used to input the image sample to be classified into the base class model and the classification models of each incremental session, obtain the classification results of the base class model and each classification model, and calculate the uncertainty value of each model; The classification result determination unit is used to select the classification result of the base class model or the classification model with the smallest uncertainty value as the classification result of the image sample to be classified.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1-7.

10. A computer storage medium, characterized in that, At least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to execute the method according to any one of claims 1-7.