Few-shot Continual Learning Method and System Based on Adaptive Adaptation Layer Selection
The adaptive layer selection method in small-sample continuous learning uses singular value decomposition to separate network weights and dynamically select layers based on adapter sensitivity ratios, enhancing classification accuracy and adaptability by minimizing interference with existing knowledge.
Patent Information
- Application Number
- CN202510007379.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing small sample continuous learning method fails to effectively balance the adaptability of new tasks with old knowledge retention at different adaptation stages, resulting in a degradation of model performance.
Adaptive adaptation layer selection method is adopted to decompose network weights into knowledge-sensitive components and redundant capacity components through singular value decomposition, dynamically evaluate and select the adapter, freeze knowledge-sensitive components, and build adapters using redundant capacity components to achieve protection of old knowledge and learning of new knowledge.
The dynamic adaptation layer selection mechanism effectively reduces interference from old knowledge, improves model classification accuracy and adaptability, maintains the stability of the model under new tasks, and reduces computing and storage overhead.
Smart Images

Figure CN119397366B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and deep learning, and particularly relates to a few-shot continual learning method and system based on adaptive adapter layer selection. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Few-shot continual learning can learn new classes with limited data while protecting existing knowledge. In practical applications, such as adaptive recommendation systems and robot navigation in dynamic environments, few-shot continual learning algorithms need to ensure that the model can effectively adapt to new information with limited resources without affecting the stability of existing knowledge. Existing methods usually use fixed network layers for fine-tuning or insert adapters at different adaptation stages, failing to consider that the ability of different network layers to retain old knowledge is dynamically changing during the continual learning process. With the continuous introduction of new tasks, the adaptability of different layers and their sensitivity to old knowledge will also change, making it difficult for manual or fixed layer selection strategies to achieve an optimal balance between adapting to new tasks and retaining old knowledge, thus affecting the performance of the few-shot continual learning model and ultimately weakening the classification accuracy of the few-shot continual learning model. Summary of the Invention
[0004] To solve the technical problems existing in the above background art, the present invention provides a few-shot continual learning method and system based on adaptive adapter layer selection, which realizes dynamic adapter selection for minimizing interference of old knowledge and ensures the classification accuracy of the few-shot continual learning model.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The first aspect of the present invention provides a few-shot continual learning method based on adaptive adapter layer selection.
[0007] A few-shot continual learning method based on adaptive adapter layer selection includes:
[0008] Based on the replay sample data and the backbone network of the continual learning model, calculate the covariance matrix of the activation values corresponding to the linear layer of the continual learning model for the replay sample data; multiply the pre-trained linear weight matrix by the covariance matrix of the activation values, and then perform singular value decomposition into a knowledge-sensitive component and a redundant capacity component;
[0009] Freeze the pre-trained linear weight matrix corresponding to the knowledge-sensitive components, multiply it by the training sample features in the current incremental adaptation training stage to obtain knowledge-sensitive features; at the same time, use the redundant capacity components to construct a learnable adapter, and adaptively determine the adaptation layer according to the pre-computed adapter sensitivity ratio of each linear layer; multiply the adapter matrix of the adaptation layer by the training sample features in the current incremental adaptation training stage to obtain redundant capacity features;
[0010] After superimposing the knowledge-sensitive features and the redundant capacity features, the output features obtained by the linear layer transformation are passed to the next layer of the continual learning model to obtain the updated pre-trained linear weight matrix;
[0011] Re-obtain the few-shot replay data, and perform singular value decomposition, adaptation layer adaptive determination, and incremental adaptation training operations in sequence based on the updated pre-trained linear weight matrix until the continual learning model reaches the set requirements and stops learning, so as to use the trained continual learning model to perform classification tasks.
[0012] As an implementation, arrange the adapter sensitivity ratios of all linear layers in ascending order, and select the K layers with the smallest adapter sensitivity ratio values for adapter construction; where K is a positive integer greater than or equal to 1.
[0013] As an implementation, the adapter sensitivity ratio is:
[0014]
[0015] where, is the adapter sensitivity ratio; is the layer singular value diagonal matrix the penultimate smallest non-zero singular value, that is, the singular value corresponding to the first component in the redundant capacity component; corresponding to the smallest non-zero singular value in
[0016] As an implementation, the expression of the covariance matrix of the activation values is:
[0017] ,
[0018] where, represents the covariance matrix; represents the activation value of the replay sample features at the current layer; represents transpose; is the number of replay samples; represents the input dimension of the pre-trained linear weight matrix.
[0019] As an implementation, the process of singular value decomposition is as follows:
[0020] ,
[0021] Among them, represents singular value decomposition; and are orthogonal matrices, is a diagonal matrix containing singular values arranged from large to small ; represents the covariance matrix; represents the pre-trained linear weight matrix; represents the rank of the pre-trained linear weight matrix; represents the left singular vector; is the th row of the matrix represents the transpose of the matrix.
[0022] As an implementation, the expression of the updated pre-trained linear weight matrix is: ;
[0023] Among them, represents the updated pre-trained linear weight matrix; represents the pre-trained linear weight matrix corresponding to the knowledge-sensitive component; and represent the adapter update parameter matrix corresponding to the redundant capacity component.
[0024] As an implementation, the training sample features in the current incremental adaptation training stage are extracted from the training sample data in the current incremental adaptation training stage; the training sample data in the current incremental adaptation training stage consists of initial training sample data and replay sample data.
[0025] The second aspect of the present invention provides a few-shot continual learning system based on adaptive adapter selection.
[0026] A few-shot continual learning system based on adaptive adapter selection includes:
[0027] A weight decomposition module, which is used to calculate the covariance matrix of the activation values corresponding to the linear layer of the continual learning model for the replay sample data based on the replay sample data and the backbone network of the continual learning model; multiply the pre-trained linear weight matrix by the covariance matrix of the activation values, and then perform singular value decomposition into knowledge-sensitive components and redundant capacity components;
[0028] An adaptation layer determination module, which is used to freeze the pre-trained linear weight matrix corresponding to the knowledge-sensitive components, multiply it by the training sample features in the current incremental adaptation training stage to obtain knowledge-sensitive features; at the same time, use the redundant capacity components to construct a learnable adapter, and adaptively determine the adaptation layer according to the pre-computed adapter sensitivity ratios of each linear layer; multiply the adapter matrix of the adaptation layer by the training sample features in the current incremental adaptation training stage to obtain redundant capacity features;
[0029] A weight update module, which is used to stack the knowledge-sensitive features and the redundant capacity features to obtain the output features obtained by the linear layer transformation and transfer them to the next layer of the continual learning model to obtain an updated pre-trained linear weight matrix;
[0030] An adaptation training module, which is used to re-obtain the few-shot replay data, and successively perform singular value decomposition, adaptation layer adaptive determination, and incremental adaptation training operations based on the updated pre-trained linear weight matrix until the continual learning model reaches the set requirements and stops learning, so as to use the trained continual learning model to perform classification tasks.
[0031] The third aspect of the present invention provides a computer-readable storage medium.
[0032] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the few-shot continual learning method based on adaptive adaptation layer selection as described above.
[0033] The fourth aspect of the present invention provides an electronic device.
[0034] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the few-shot continual learning method based on adaptive adaptation layer selection as described above.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] (1) Dynamic adaptation layer selection mechanism: Different from the traditional method of manually selecting the adaptation layer, the present invention dynamically evaluates the ability of each network layer to retain old knowledge and adapt to new tasks by introducing a quantization index, the adapter sensitivity ratio (ASR). By sorting and selecting the K layers with the lowest ASR values as the adaptation layers, the present invention realizes dynamic adapter selection for minimizing the interference of old knowledge.
[0037] (2) Knowledge-sensitive decomposition and redundant capacity modeling: Singular Value Decomposition (SVD) is used to decompose the network weights into knowledge-sensitive components and redundant capacity components, which are used to retain existing knowledge and adapt to new tasks respectively. The knowledge-sensitive components are frozen to protect old knowledge, while the redundant capacity components are constructed as adapters and fine-tuned to learn new knowledge, thus effectively solving the problem of the balance between adaptability and stability in traditional methods when freezing weights or fully fine-tuning.
[0038] (3) Dynamic evaluation and continuous adjustment: The present invention can perform real-time dynamic evaluation on network layers at each incremental stage, recalculate the ASR based on the latest replay data and select the adapted layer to adapt to the task requirements at different stages, overcoming the problem of the lack of dynamic adjustment ability in the existing fixed layer selection strategy.
[0039] (4) Low computational and storage overhead: Compared with the data replay method that requires a large amount of storage space or the dynamic expansion method that introduces additional model parameters, the present invention utilizes the redundant capacity in the existing network structure, avoiding modifying the model architecture or significantly increasing the inference overhead, and having higher efficiency and practicality.
[0040] (5) Adapted layer optimization and knowledge integration: After adaptation, the fine-tuned redundant capacity components and the frozen knowledge-sensitive components are recombined to restore the original structure of the weight matrix, and new knowledge is integrated into the knowledge-sensitive components by continuously updating the covariance matrix to ensure that the model can better adapt to new tasks in future stages.
[0041] The advantages of additional aspects of the present invention will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings
[0042] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0043] Figure 1 is a flowchart of a few-shot continual learning method based on adaptive adapted layer selection according to an embodiment of the present invention;
[0044] Figure 2 is a schematic structural diagram of a few-shot continual learning system based on adaptive adapted layer selection according to an embodiment of the present invention. Detailed Embodiments
[0045] The present invention will be further described below in conjunction with the drawings and embodiments.
[0046] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.
[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0048] Term Explanation:
[0049] Few-shot Continual Learning:
[0050] Aims to learn a model that can continuously learn knowledge from a series of new classes using only a small number of labeled training samples while retaining the knowledge obtained from learning previous classes. The data for each incremental adaptation stage can be represented as . In each training stage t, the model receives the training set , which contains the input samples and their corresponding labels . The initial training stage is also referred to as the base training stage, and its label set has a sufficient number of labels and samples to enable the model to obtain good basic representation capabilities and lay a foundation for subsequent continual learning. In subsequent training stages , the model is gradually introduced to new classes, with only a small number of labeled samples for each class, and there is no overlap between the class sets of different training stages (such as stage t and stage t') . During the training process of the new stage, only limited replay samples from the learned classes (usually only one replay sample per class) can be accessed to retain the learned knowledge, and the complete data from the previous training stage cannot be directly obtained. In the model evaluation stage after the training stage t, the model needs to be tested on all learned classes to evaluate the model's ability to learn new classes and retain the learned classes.
[0051] Embodiment 1
[0052] According to Figure 1 shown, this embodiment provides a few-shot continual learning method based on adaptive adaptation layer selection, including:
[0053] Step 1: Based on the replay sample data and the backbone network of the continual learning model, calculate the covariance matrix of the activation values corresponding to the replay sample data in the linear layer of the continual learning model; multiply the pre-trained linear weight matrix by the covariance matrix of the activation values, and then perform singular value decomposition into knowledge-sensitive components and redundant capacity components;
[0054] Step 2: Freeze the pre-trained linear weight matrix corresponding to the knowledge-sensitive components, multiply it by the training sample features in the current incremental adaptation training stage to obtain knowledge-sensitive features; at the same time, use the redundant capacity components to construct a learnable adapter, and adaptively determine the adaptation layer according to the pre-calculated adapter sensitivity ratios of each linear layer; multiply the adapter matrix of the adaptation layer by the training sample features in the current incremental adaptation training stage to obtain redundant capacity features;
[0055] Step 3: After superimposing the knowledge-sensitive features and the redundant capacity features, obtain the output features transformed by the linear layer and pass them to the next layer of the continual learning model to obtain the updated pre-trained linear weight matrix;
[0056] Step 4: Re-obtain the few-shot replay data, and successively perform singular value decomposition, adaptation layer adaptive determination, and incremental adaptation training operations based on the updated pre-trained linear weight matrix (i.e., repeat Steps 1 to 3) until the continual learning model reaches the set requirements and stops learning, so as to use the trained continual learning model to perform classification tasks.
[0057] To achieve the above object, in Step 1, at the beginning of each incremental training stage before starting, collect a small replay sample data set , where represents the union; each old class contains only one sample. Input the replay samples into the backbone network of the model to calculate the activation values of the intermediate layer of the model, and calculate the covariance matrix before each linear layer :
[0058] ,
[0059] where represents the covariance matrix; represents the activation value of the replay sample features at the current layer; represents transpose; is the number of replay samples; represents the input dimension of the pre-trained linear weight matrix.
[0060] Among them, the process of performing singular value decomposition on the product of the covariance matrix of each activation value and the linear weight matrix is:
[0061] ,
[0062] Among them, represents singular value decomposition; and are orthogonal matrices, is a diagonal matrix containing singular values arranged from large to small ; represents the covariance matrix; represents the pre-trained linear weight matrix; represents the rank of the pre-trained linear weight matrix; represents the left singular vector; is the th row of the matrix represents the transpose of the matrix.
[0063] In the specific implementation process, the first R - r singular vectors with larger singular values are used as knowledge-sensitive components; the last r singular vectors with smaller corresponding singular values are used as redundant capacity components; where R represents the rank of the pre-trained linear weight matrix, which is the number of non-zero singular values, , represents the input dimension of the pre-trained linear weight matrix; represents the output dimension of the pre-trained linear weight matrix.
[0064] Generally, r is set to 64; the selection basis of the r value can be determined according to the actual application or experimental results, usually determined by cross-validation or model performance evaluation.
[0065] Among them, the knowledge-sensitive component and the redundant capacity component are complementary. The redundant capacity component can minimize the interference to the existing capabilities to ensure the stability of the old categories and the plasticity of the new tasks.
[0066] In the specific implementation process of step 2, after the weight decomposition, the weights of the model are divided into two parts: "knowledge-sensitive components" and "redundant capacity components". The knowledge-sensitive components are used to retain the existing knowledge, while the redundant capacity components provide a flexible space for learning new knowledge. Although all network layers contain redundant capacity components, not every layer has the same ability to protect the existing knowledge, nor does every layer have the same redundant capacity to adapt to new tasks. When fine-tuning the adapter, if the adapter is assigned to those layers that are more sensitive to the existing knowledge, it may damage the original knowledge structure and lead to catastrophic forgetting. Therefore, it is necessary to accurately select the most suitable layer for adapting to the new task before each incremental adaptation stage, while ensuring the minimum interference to the existing knowledge.
[0067] In an embodiment of the present invention, an adaptive adapter layer selection strategy is proposed, that is, the adapter layer is adaptively determined according to the pre-computed Adapter Sensitivity Ratio (ASR) of each linear layer.
[0068] The adapter sensitivity ratio quantifies the sensitivity of the separated redundant capacity components in each layer to the existing knowledge. A lower ASR indicates that the redundant capacity components have singular values closer to the minimum value, and the adapters constructed with these components are less likely to interfere with the existing knowledge. On the contrary, a higher ASR indicates that the adapters contain some important components that make non-negligible contributions to the existing knowledge. Therefore, the layers with lower ASR should be preferentially used for adapting new tasks. Before the start of each new incremental adaptation training, the adaptive layer selection is performed through the following steps:
[0069] (1) Obtain the singular value matrix: At the t-th adaptation stage, perform knowledge-preserving weight decomposition on each linear layer in the backbone network to obtain the singular value diagonal matrix of each layer , for example, the singular value diagonal matrix obtained after performing knowledge-preserving singular value decomposition on the linear weight matrix of the -th layer is ;
[0070] (2) Adapter sensitivity ratio calculation: Calculate the adapter sensitivity ratio of each linear layer according to the singular values of the calculated redundant capacity components :
[0071]
[0072] where is the -th smallest non-zero singular value of the singular value diagonal matrix , that is, the singular value corresponding to the first component in the redundant capacity components; corresponds to the smallest non-zero singular value in
[0073] .
[0073] (3) Adapter sensitivity ratio sorting: After arranging the N ASR values in descending order, we get:
[0074] , , ...,
[0075] where is the position of the layer where the minimum ASR value among the N ASR values is located, is the corresponding minimum ASR value; is the position of the layer where the second smallest ASR value among the N ASR values is located. is the corresponding second smallest ASR value; is the position of the layer where the largest ASR value among the N ASR values is located, is the corresponding largest ASR value.
[0076] (4) Select the K layers with the smallest sensitivity as the adaptation layers: Select the last K layers with the smallest ASR values to construct an adapter based on knowledge protection decomposition. The positions of the selected layers are as follows:
[0077]
[0078] where K is the number of adaptation layers set artificially, used to control the number of learnable parameters, and is default set to 6; is the set of the K layers with the least impact on the preservation of existing knowledge selected by the adaptive layer selection mechanism proposed by the present invention before the start of the t-th adaptation stage; is the position serial number of the layer with the smallest ASR value; is the serial number of the position of the layer with the second smallest ASR value; is the position requirement of the layer with the K-th smallest ASR value. The adapter constructed by these K layers is used for parameter update in the t-th stage, and the remaining adapters are not adapted.
[0079] The redundant capacity component is used to construct two learnable low-rank matrices and as adapters:
[0080]
[0081]
[0082]
[0083]
[0084]
[0085] Among them, represents the pre-trained linear weight matrix corresponding to the knowledge-sensitive component; represents the pre-trained linear weight matrix; and are orthogonal matrices; represents the inverse of the covariance matrix; represents the left singular vector; represents the singular value; is the th row of the matrix denotes the transpose of a matrix; denotes the last column, and respectively denote the adapter matrices; During the adaptation process, by only updating and while freezing to maintain the learned knowledge.
[0086] In step 3, during the incremental training process at the t-th stage, each training stage includes a new set of training samples and corresponding labels. The input data can be images, text, or other forms of samples. These data first pass through the input layer of the backbone network, are passed layer by layer to the linear layer, and finally are adjusted by the adapter. For each input sample, the features of the i-th training sample are processed by the linear layer in the network. The calculation method of the current linear layer is as follows:
[0087]
[0088] where denotes the pre-trained linear weight matrix corresponding to the knowledge-sensitive component; B and A are a set of learnable adapter matrices; is the input feature of the current linear layer; is the output feature obtained after the transformation of the current linear layer and is passed to the next layer of the network.
[0089] In this process, the role of the adapter is to fine-tune the weights of the network. Especially during the training process of adding new classes, the adapter will perform adaptive fine-tuning on the basis of the original model. The adapter effectively integrates new knowledge into the original classification ability by calculating the combination of the input features and the new weight matrix.
[0090] The loss function is usually used to measure the difference between the model output and the true label. In incremental training, the calculation of the loss function usually depends on the output of the network and the true label (i.e., the actual label of the input sample). Suppose we use the cross-entropy loss function as the loss function in common classification tasks, and its calculation formula is:
[0091]
[0092] where denotes the loss value of the i-th training sample; denotes the true label of the i-th sample on class c (usually 0 or 1, indicating whether it belongs to this class); denotes the predicted probability of the i-th sample on class c, usually obtained through the softmax layer of the model.
[0093] After calculating the loss, the backpropagation algorithm is used to calculate the gradients of each model parameter. The gradient represents the rate of change of the loss function with respect to the adapter parameters. Through backpropagation, the gradients of the loss with respect to the weight matrices A and B are calculated, and the model parameters are updated based on these gradients to minimize the loss.
[0094] The expression for the updated pre-trained linear weight matrix is:
[0095]
[0096] Where, represents the updated pre-trained linear weight matrix; W represents the pre-trained linear weight matrix corresponding to the knowledge-sensitive component; and represent the adapter update parameter matrices corresponding to the redundant capacity components.
[0097] Specifically, in step 4, the covariance matrix is recalculated before the start of each new adaptation phase and the above decomposition process is repeated, so as to integrate the newly learned capabilities into the knowledge-sensitive components, enabling the method to adapt to the evolving model capabilities in few-shot continual learning. The replay dataset is updated in the th adaptation phase to include all previously learned classes:
[0098]
[0099] Then, the knowledge-protecting weight decomposition of the N linear weight layers is performed again using the updated replay data. After decomposition, the new adaptation layers obtained after performing the above adaptive layer selection mechanism are as follows:
[0100] 。
[0101] Where is the set of K layers with the least impact on the retention of existing knowledge selected by the adaptive layer selection mechanism proposed by the present invention before the start of the (t + 1)-th adaptation phase; is the position index of the layer with the smallest ASR value calculated after performing the knowledge-protecting decomposition using the latest replay dataset ; is the position index of the layer with the second smallest ASR value calculated after performing the knowledge-protecting decomposition on the latest replay dataset ; is the position index of the layer with the K-th smallest ASR value calculated after performing the knowledge-protecting decomposition on the latest replay dataset . The adapters constructed by these K layers are used for parameter update in the (t + 1)-th phase, and the remaining adapters are not updated.
[0102] In summary, the above framework can effectively protect the learned knowledge and improve the adaptation performance to new tasks without changing the model architecture or increasing the inference overhead by applying continuous knowledge protection decomposition to the layers adaptively selected in the backbone model with the least impact on existing knowledge.
[0103] The continuous knowledge protection decomposition method proposed by the present invention has the following remarkable beneficial effects compared with the prior art:
[0104] (1) Dynamic adaptation layer selection mechanism: Different from the traditional method of manually selecting adaptation layers, the present invention dynamically evaluates the ability of each network layer to retain old knowledge and adapt to new tasks by introducing a quantization index, the adapter sensitivity ratio (ASR). By sorting and selecting the K layers with the lowest ASR values as the adaptation layers, the present invention realizes the dynamic adapter selection for minimizing the interference of old knowledge.
[0105] (2) Knowledge-sensitive decomposition and redundant capacity modeling: The network weights are decomposed into knowledge-sensitive components and redundant capacity components by using singular value decomposition (SVD), which are used to retain existing knowledge and adapt to new tasks respectively. The knowledge-sensitive components are frozen to protect old knowledge, while the redundant capacity components are constructed as adapters and fine-tuned to learn new knowledge, thus effectively solving the problem of the balance between adaptability and stability in the traditional methods when freezing weights or fully fine-tuning.
[0106] (3) Dynamic evaluation and continuous adjustment: The present invention can perform real-time dynamic evaluation on network layers at each incremental stage, recalculate the ASR according to the latest replay data and select the adaptation layers to adapt to the task requirements at different stages, overcoming the problem of the lack of dynamic adjustment ability in the existing fixed layer selection strategy.
[0107] (4) Low computational and storage overhead: Compared with the data replay method that requires a large amount of storage space or the dynamic expansion method that introduces additional model parameters, the present invention avoids modifying the model architecture or significantly increasing the inference overhead by utilizing the redundant capacity in the existing network structure, and has higher efficiency and practicality.
[0108] (5) Adaptation layer optimization and knowledge integration: After the adaptation is completed, the fine-tuned redundant capacity components and the frozen knowledge-sensitive components are recombined to restore the original structure of the weight matrix, and new knowledge is integrated into the knowledge-sensitive components by continuously updating the covariance matrix to ensure that the model can better adapt to new tasks in future stages.
[0109] In summary, by introducing an adaptive adaptation layer selection mechanism and a weight decomposition strategy, the present invention effectively balances the learning of new knowledge and the retention of old knowledge without modifying the model architecture, significantly improving the adaptability and generalization ability of few-shot continual learning. This innovative method is significantly superior to the prior art in terms of model efficiency, knowledge protection, and dynamic adaptability.
[0110] Table 1 Comparison of the present invention with other methods on miniImageNet
[0111]
[0112] Wherein:
[0113] DSN: It is a few-shot continual learning algorithm based on the Dynamic Support Network (DSN), which uses the way of expanding the nodes of the dynamic support network to support the adaptive update network of the feature space.
[0114] Data-free: It is a few-shot continual learning algorithm based on Data Free Replay. It synthesizes the learned data through a generator to keep the learned knowledge from being forgotten without accessing the real data.
[0115] MetaFSCIL: It is a few-shot class-incremental learning algorithm (Few-shot Class-incremental Learning, FSCIL) based on meta learning, which uses meta-goals to learn so as to quickly adapt to new tasks without forgetting the learned knowledge.
[0116] FeSSSS: It is a method based on the FEw-shot Self-Supervised Syetem (FeSSSS). It uses the progress brought by self-supervised learning to correct overfitting and catastrophic forgetting, and trains lightweight feature fusion and classifiers on a series of features that appear in supervised and self-supervised models.
[0117] C-FSCIL: It is a Constrained Few-shot Class-Incremental Learning (C-FSCIL) algorithm, which can trade off between the accuracy of learning new classes and the computational memory cost.
[0118] LIMIT: It is a few-shot continual learning algorithm through Learning Multi-phase Incremental Tasks (LIMIT). By utilizing the base class data to construct pseudo few-shot continual learning tasks, it builds a generalizable feature space for unseen tasks.
[0119] FACT: It is a few-shot continual learning algorithm based on Forward Compatible Training (FACT). By allocating virtual cluster centers to squeeze the embedding space of known classes, it provides a feature space for the embeddings of new classes.
[0120] TEEN: It is a few-shot continual learning algorithm based on the Training-free calibration strategy. By fusing new cluster centers (i.e., the average features of classes) with weighted base prototypes, it enhances the discriminability of new classes.
[0121] ALICE: It is a few-shot incremental learning algorithm based on Augmented Angular Loss Incremental Classification (ALICE). It uses angular penalty loss to replace the commonly used cross-entropy loss to obtain well-clustered features.
[0122] CABD: It is a few-shot continual learning algorithm based on Class-Aware Bilateral Distillation (CABD), which is used to adaptively learn knowledge from two complementary teachers.
[0123] NC-FSCIL: It is a few-shot class incremental learning framework inspired by neural collapse, Neural Collapse inspired feature-classifier alignment for Few-Shot Class Incremental Learning (NC-FSCIL). By using a pre-allocated classifier, it alleviates the misalignment between the features of old classes and the classifier faced in few-shot continual learning.
[0124] Mamba-FSCIL: It is a few-shot class incremental learning algorithm (Few-shot Class-incremental Learning, FSCIL) based on the selective state space model (i.e., the Mamba model). It designs a projector for the dual selective state space model, which adaptively adjusts its projection parameters according to the input context to adapt to new classes while reducing interference to existing classes.
[0125] CEC+:It is a few-shot class-incremental learning algorithm based on Continually Evolved Classifiers (CEC). It updates classification weights and test features through a knowledge-guided relationship refinement module and designs a pseudo-incremental relationship refinement learning method to mine global and local concepts using a novel concept mining strategy.
[0126] KANet:It is a few-shot class-incremental learning algorithm based on the Knowledge Adaptation Network. It summarizes data-specific knowledge from training data and integrates it into the general representation.
[0127] CPE-CLIP:It is a few-shot class-incremental learning algorithm that performs Continual Parameter-Efficient finetuning (CPE) on the CLIP model (based on Contrastive Language-Image Pre-training, CLIP), leveraging the large amount of knowledge obtained by CLIP in large-scale pre-training and its effectiveness in promoting new concepts.
[0128] PL-FSCIL:It is a few-shot class-incremental learning method based on Prompt Learning for Few-shot Class-incremental Learning. This method combines a visual promptor with a pre-trained vision Transformer to effectively address the challenges faced in few-shot continual learning.
[0129] PriViLege:It is a pre-trained vision and language model based on prompting functions and knowledge distillation (Pre-trained Vision and Language transformers with prompting functions and knowledge distillation). It effectively addresses the challenges of catastrophic forgetting and overfitting in large models through new pre-trained knowledge adaptation and two loss functions.
[0130] As can be seen from Table 1, the present invention has the following advantages:
[0131] (1) Higher average accuracy: The adaptive adaptation layer selection method proposed in the present invention performs excellently in multiple incremental learning continuous training stages. Especially on the CIFAR-100 dataset, it demonstrates a higher accuracy than existing methods. On the CIFAR-100 dataset, the average accuracy of the adaptive adaptation layer selection method proposed in the present invention is 88.62%, far exceeding existing methods such as PL-FSCIL (72.66%) and PriViLege (88.41%).
[0132] (2) Stronger knowledge retention ability: In the incremental learning task of gradually introducing new categories, the present invention can effectively adapt to new categories and retain the memory of previous knowledge. For example, in the last continuous learning stage of CIFAR-100, the accuracy of the continuous decomposition adaptation method proposed in the present invention is 86.22%, significantly superior to other methods, proving that this method has good adaptability when new categories are introduced and can effectively reduce catastrophic forgetting.
[0133] The adaptive adaptation layer selection method of the present invention shows superior performance in incremental learning and few-shot classification tasks, can effectively cope with the gradual introduction of new categories, and surpasses existing state-of-the-art technologies on the CIFAR-100 dataset.
[0134] The present invention has broad application prospects, especially in products and systems that require efficient processing of incremental learning tasks. For example, the following application fields:
[0135] Intelligent vision systems: In visual perception systems such as intelligent cameras, autonomous vehicles, and drones, the present invention can improve the continuous learning ability of the system under limited labeled data, avoid catastrophic forgetting, retain the memory of existing knowledge, and optimize the accuracy of tasks such as object detection and behavior analysis.
[0136] Robots and autonomous systems: In the fields of robot navigation, autonomous drones, intelligent manufacturing, and automated production, the present invention improves the flexibility and stability of the system by helping the system quickly adapt to new tasks in a changing environment while retaining the ability of existing tasks, and is suitable for long-term operation and multi-task scenarios.
[0137] Augmented reality and mobile applications: The present invention can be integrated into augmented reality applications, smartphone camera applications, social media filters, etc. to help improve image processing and computer vision functions. Through continuous learning and adaptation, the present invention can enhance the user experience and maintain efficient and accurate visual effects in a dynamic environment.
[0138] The few-shot continuous learning method based on adaptive adaptation layer selection of the present invention is widely used in the following products:
[0139] Intelligent Vision Products: The present invention can be applied to devices such as intelligent cameras, autonomous vehicles, and security monitoring systems. By continuously learning, it enhances visual perception capabilities and supports rapid learning of new categories and effective retention of existing knowledge without a large amount of labeled data.
[0140] Robots and Automation Equipment: The present invention is applicable to autonomous systems such as robot navigation, autonomous drones, and automated production. It helps improve visual perception, task adaptation, and decision-making capabilities, enabling the system to maintain efficient learning and stable performance in a dynamic environment.
[0141] Augmented Reality and Mobile Applications: The present invention can be integrated into fields such as augmented reality systems, smartphone applications, and social media filters. It helps optimize image processing, object recognition, and content generation functions, enhances the user experience, and improves the system's ability to adapt to environmental changes.
[0142] Embodiment 2
[0143] As Figure 2 shown, in this embodiment, a few-shot continual learning system based on adaptive adapter layer selection is provided, including:
[0144] A weight decomposition module 201, which is used to calculate the covariance matrix of the activation values corresponding to the replay sample data in the linear layer of the continual learning model based on the replay sample data and the backbone network of the continual learning model; multiply the pre-trained linear weight matrix by the covariance matrix of the activation values, and then perform singular value decomposition into a knowledge-sensitive component and a redundant capacity component;
[0145] An adapter layer determination module 202, which is used to freeze the pre-trained linear weight matrix corresponding to the knowledge-sensitive component, multiply it by the training sample features in the current incremental adaptation training stage to obtain knowledge-sensitive features; at the same time, use the redundant capacity component to construct a learnable adapter, and adaptively determine the adapter layer according to the pre-calculated adapter sensitivity ratio of each linear layer; multiply the adapter matrix of the adapter layer by the training sample features in the current incremental adaptation training stage to obtain redundant capacity features;
[0146] A weight update module 203, which is used to stack the knowledge-sensitive features and the redundant capacity features to obtain the output features transformed by the linear layer and transfer them to the next layer of the continual learning model to obtain an updated pre-trained linear weight matrix;
[0147] An adaptation training module 204, which is used to re-obtain the few-shot replay data, and sequentially perform singular value decomposition, adapter layer adaptive determination, and incremental adaptation training operations based on the updated pre-trained linear weight matrix until the continual learning model reaches the set requirements and stops learning, so as to use the trained continual learning model to perform classification tasks.
[0148] It should be noted here that each module in this embodiment corresponds one by one to each step in the first embodiment, and the specific implementation process is the same, so it will not be elaborated here.
[0149] Embodiment Three
[0150] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the few-shot continual learning method based on adaptive adapter selection as described above.
[0151] Embodiment Four
[0152] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the few-shot continual learning method based on adaptive adapter selection as described above.
[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple blocks.
[0154] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A few-shot continual learning method based on adaptive adaptation layer selection, characterized in that Including: Based on the replay picture sample data and the backbone network of the continual learning model, calculate the covariance matrix of the activation values corresponding to the replay picture sample data in the linear layer of the continual learning model; Multiply the pre-trained linear weight matrix by the covariance matrix of the activation values, and then perform singular value decomposition into a knowledge-sensitive component and a redundant capacity component; Freeze the pre-trained linear weight matrix corresponding to the knowledge-sensitive component, multiply it by the picture training sample features in the current incremental adaptation training stage to obtain knowledge-sensitive features; at the same time, use the redundant capacity component to construct a learnable adapter, and adaptively determine the adaptation layer according to the pre-computed adapter sensitivity ratio of each linear layer; multiply the adapter matrix of the adaptation layer by the picture training sample features in the current incremental adaptation training stage to obtain redundant capacity features; After superimposing the knowledge-sensitive features and the redundant capacity features, obtain the output features transformed by the linear layer and transfer them to the next layer of the continual learning model to obtain the updated pre-trained linear weight matrix; Re-obtain the small picture sample replay data, and successively perform singular value decomposition, adaptation layer adaptive determination and incremental adaptation training operations based on the updated pre-trained linear weight matrix until the continual learning model stops learning when the set requirements are met, so as to use the trained continual learning model to perform picture classification tasks; Arrange the adapter sensitivity ratios of all linear layers in ascending order, and select the top K layers with the smallest adapter sensitivity ratio values for adapter construction; where K is a positive integer greater than or equal to 1; The small-sample replay dataset of pictures is updated in the th adaptation stage to: ; Among them, is the latest small sample replay dataset of pictures; is the input sample; is the label corresponding to the input sample; is the union set; is the category set in the i-th training stage; represents the union set of the category sets from the 0-th to the t-th stage in the training stage; Use the updated small picture sample replay data to perform knowledge-protected weight decomposition on N linear weight layers; after decomposition, execute the adaptive layer selection mechanism; obtain the new adaptation layer: ; Among them, is the set of K layers with the least impact on the existing knowledge selected by the adaptive layer selection mechanism before the start of the (t + 1)-th adaptation stage; is the position index of the layer with the smallest ASR value calculated after performing knowledge protection decomposition using the latest small sample replay dataset of images ; is the position index of the layer with the second smallest ASR value calculated after performing knowledge protection decomposition on the latest small sample replay dataset of images ; is the position index of the layer with the K-th smallest ASR value calculated after performing knowledge protection decomposition on the latest small sample replay dataset of images .
2. The few-shot continual learning method based on adaptive adaptation layer selection according to claim 1, wherein The adapter sensitivity ratio is: Among them, is the ratio of adapter sensitivities; is the layer singular value diagonal matrix the penultimate small non-zero singular value, that is, the singular value corresponding to the first component in the redundant capacity component; corresponds to the smallest non-zero singular value in 3. The few-shot continual learning method based on adaptive adaptation layer selection according to claim 1, wherein The expression of the covariance matrix of the activation values is: , Among them, represents the covariance matrix; represents the activation value of the replay picture sample feature at the current layer; represents transpose; is the number of replay picture samples; represents the input dimension of the pre-trained linear weight matrix.
4. The few-shot continual learning method based on adaptive adaptation layer selection according to claim 1, characterized in that, The process of singular value decomposition is: , Among them, denotes singular value decomposition; and are orthogonal matrices, is a diagonal matrix containing singular values arranged from large to small ; denotes the covariance matrix; denotes the pre-trained linear weight matrix; denotes the rank of the pre-trained linear weight matrix; denotes the left singular vector; is the th row of the matrix denotes the transpose of the matrix.
5. The few-shot continual learning method based on adaptive adaptation layer selection according to claim 1, wherein The expression for the updated pre-trained linear weight matrix is as follows: ; Among them, represents the updated pre-trained linear weight matrix; represents the pre-trained linear weight matrix corresponding to the knowledge-sensitive component; and represents the adapter update parameter matrix corresponding to the redundant capacity component.
6. The few-shot continual learning method based on adaptive adaptation layer selection according to claim 1, wherein The picture training sample features in the current incremental adaptation training stage are extracted from the picture training sample data in the current incremental adaptation training stage; the picture training sample data in the current incremental adaptation training stage is composed of the initial picture training sample data and the replay picture sample data.
7. A few-shot continual learning system based on adaptive adaptation layer selection, characterized in that, Including: A weight decomposition module, which is used to calculate the covariance matrix of the activation values corresponding to the replay picture sample data in the linear layer of the continual learning model based on the replay picture sample data and the backbone network of the continual learning model; multiply the pre-trained linear weight matrix by the covariance matrix of the activation values, and then perform singular value decomposition into a knowledge-sensitive component and a redundant capacity component; An adaptation layer determination module, which is used to freeze the pre-trained linear weight matrix corresponding to the knowledge-sensitive component, multiply it by the picture training sample features in the current incremental adaptation training stage to obtain knowledge-sensitive features; at the same time, use the redundant capacity component to construct a learnable adapter, and adaptively determine the adaptation layer according to the pre-computed adapter sensitivity ratio of each linear layer; multiply the adapter matrix of the adaptation layer by the picture training sample features in the current incremental adaptation training stage to obtain redundant capacity features; A weight update module, which is used to superimpose the knowledge-sensitive features and the redundant capacity features, obtain the output features obtained by linear layer transformation, and transmit them to the next layer of the continual learning model to obtain an updated pre-trained linear weight matrix; An adaptation training module, which is used to re-obtain the small sample replay data of pictures, and perform singular value decomposition, adaptation layer self-adaptation determination, and incremental adaptation training operations in sequence based on the updated pre-trained linear weight matrix until the continual learning model meets the set requirements and stops learning, so as to use the trained continual learning model to perform picture classification tasks; Arrange the adapter sensitivity ratios of all linear layers in ascending order, and select the top K layers with the smallest adapter sensitivity ratio values for adapter construction; where K is a positive integer greater than or equal to 1; The small-sample replay dataset of pictures is updated in the th adaptation phase as follows: ; Among them, is the latest small sample replay dataset of pictures; is the input sample; is the label corresponding to the input sample; is the union set; is the class set in the i-th training stage; represents the union of the class sets from the 0-th to the t-th stage in the training stage; Use the updated small sample replay data of pictures to perform knowledge protection weight decomposition on N linear weight layers; after decomposition, execute the adaptive layer selection mechanism; obtain a new adaptation layer: ; Among them, is the set of K layers with the least impact on the existing knowledge selected by the adaptive layer selection mechanism before the start of the (t + 1)-th adaptation stage; is the position serial number of the layer with the smallest ASR value calculated after performing knowledge protection decomposition using the latest small sample replay dataset of pictures ; is the position serial number of the layer with the second smallest reciprocal ASR value calculated after performing knowledge protection decomposition on the latest small sample replay dataset of pictures ; is the position serial number of the layer with the K-th smallest reciprocal ASR value calculated after performing knowledge protection decomposition on the latest small sample replay dataset of pictures .
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the few-shot continual learning method based on adaptive adaptation layer selection as described in any one of claims 1-6.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the steps in the few-shot continual learning method based on adaptive adaptation layer selection as described in any one of claims 1-6.
Citation Information
Patent Citations
Image classification continuous learning method based on information bottleneck
CN118313438A
Persistent evolution learning method for prompting fine tuning in null space
CN118657989A