Continuous learning method based on adapter enhanced semantic prompt

By introducing adapter-enhanced semantic cues in the continuous learning method, combining language and visual information for interactive processing, and using a new cosine contrast loss function, the problem of catastrophic forgetting and generalization capabilities is solved, and more efficient and stable multi-task learning and long-sequence task processing is achieved.

CN120088539APending Publication Date: 2025-06-03BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510095335.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing continuous learning methods tend to lead to catastrophic forgetting when dealing with new tasks, and visual information-based methods are difficult to obtain knowledge of strong generalization capabilities.

Method used

A continuous learning method based on adapter-enhanced semantic prompts is proposed. By inputting language prompts and image information together into the backbone network, the self-attention module in ViT is used to process the interaction between language and images, and fine-tuning this information through the adapter. At the same time, a new cosine contrast loss function is designed to improve the output quality of the model.

Benefits of technology

It effectively solves the problem of catastrophic forgetting, improves the generalization ability and output quality of the model, and significantly improves the performance on multi-task learning and long-sequence tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088539A_ABST
    Figure CN120088539A_ABST
Patent Text Reader

Abstract

The invention discloses a continuous learning method for enhancing semantic prompt based on an adapter, which integrates a prompt fine tuning technology and an adapter fine tuning technology in the neighborhood of a fine tuning pre-training model so as to realize effective multi-modal feature adaptation. According to the method, semantic prompts generated based on a large language model are designed to enhance the generalization ability of visual features, and semantic information and visual information are effectively fused by using an adapter, so that a continuous learning task is helped to learn more adaptive features. In addition, the invention further develops a novel matching method which is used for selecting the best matching task prompt, so that the final classification accuracy of the image is indirectly improved. A wide range of experiments carried out on three continuous learning data sets prove that compared with several most advanced methods in the prior art, the method has better image classification and anti-forgetting capabilities in continuous learning tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of pattern recognition, deep learning, and continuous learning and model prompt fine-tuning, and particularly to a continuous learning method based on adapter-enhanced semantic prompts. Background Art

[0002] Continuous learning technology, as a machine learning means to cope with the continuous change of data streams, enables the model to continuously learn and adapt to new environments. Compared with traditional machine learning models, the latter are usually trained on static data sets. Once the training is completed, their knowledge bases are solidified and it is difficult to adapt to new data or tasks. The emergence of continuous learning technology aims to overcome challenges such as data privacy and memory resource limitations, allowing the model to continuously learn new knowledge on the basis of accumulating existing knowledge, rather than having to relearn every time a new task is encountered.

[0003] However, a major problem faced by continuous learning is catastrophic forgetting. This phenomenon is manifested as that when learning a new task, the update of the model weights may cause new knowledge to overwrite old knowledge, thereby affecting the performance when processing old tasks. To solve this problem, some methods propose to use knowledge distillation technology to preserve old knowledge. However, due to the lack of data for old tasks, these methods have their limitations. Therefore, some methods propose to retain representative samples in old tasks and use them to recall old knowledge. However, retaining samples may cause an excessive memory burden and pose a risk of privacy leakage. In addition, some other methods attempt to add new network branches for each new task, but this increases the size of the model, resulting in a slower inference speed and lower efficiency.

[0004] Recently, the progress of natural language processing technology, especially the development of pre-trained language models such as BERT and GPT, has provided new inspiration for continuous learning. These models can effectively understand and generate natural language, making it possible to guide continuous learning using language prompts. As a highly abstract and expressive carrier, language can not only provide high-level concepts and guidance for tasks, helping the model better summarize the learning content, but also explicitly express the relationships between different tasks, promoting the model to recognize and utilize these relationships. Therefore, in recent years, some prompt-based continuous methods have emerged, which realize using the knowledge pre-trained by the model and fine-tuning the model with fewer parameters to adapt to continuous learning tasks. These methods extract visual representations by learning a prompt library and retain the knowledge of training classes. They use visual prompts to adjust image features and keep the backbone network parameters frozen. However, layers with different depths in the backbone network have different feature refinement and learning capabilities. Ignoring this structural feature may lead to limited performance of prompt-based methods.

[0005] In addition, existing prompt-based continual learning methods solely rely on visual information, which may make it difficult for the model to acquire knowledge with strong generalization ability, especially when the training data is limited and the sample diversity is insufficient. In contrast, the class semantic information extracted by large language models is more generalizable and can adapt to various tasks. For example, models such as BERT and CLIP align the corresponding semantic information related to the class label of samples in the same category, thus creating a more general feature space. This alignment enables the model to remain robust when encountering unseen samples of the same category because the model learns to focus on common semantic features rather than just relying on specific visual features. Summary of the Invention

[0006] Aiming at the deficiencies of existing methods, the present invention proposes an innovative continual learning method that combines the prompt and adapter techniques. By inputting language prompts and image information into the backbone network together and using the self-attention module in ViT to process the interaction between language and image, and further fine-tuning this information through the adapter. In addition, to specifically constrain the language features, the present invention designs a new cosine contrast loss to improve the output quality of the model. After reexamining the initialization method of language prompts, the present invention selects the untrainable features generated by the text encoder as the initialization instead of using the traditional random initialization method to enhance the model's response to language guidance.

[0007] To address the challenges brought by continual learning, the present invention proposes a novel Adapter-Enhanced Semantic Prompting framework (AESP). The framework of the present invention includes two types of prompts: randomly initialized visual prompts and semantic prompts generated by the text encoder. Since each task requires specific prompts, selecting the most relevant task-specific prompts is crucial for effective learning. For this purpose, the present invention designs a new fusion query-key matching mechanism that combines multiple matching methods and adopts a voting mechanism, greatly improving the accuracy of prompt selection.

[0008] Once the relevant prompts are selected, they are added to the input of the network. Given the significant differences in the manifestation forms and feature spaces between visual and semantic features, the present invention integrates the adapter into each layer of the ViT network through residual connections to effectively fuse these two modalities. Therefore, the final output characteristics of the model contain visual and semantic information. To enable the model to learn general knowledge from semantic prompts, the present invention designs a novel cosine contrast loss function that can promote the alignment of semantic prompts with the semantic space of specific categories.

[0009] The technical solution steps of the present invention are as follows:

[0010] Step 1: Model initialization;

[0011] S11 Model loading: Load the pre-trained ViT and large language model BERT weights into the corresponding network structures and freeze them. Initialize the adapters and insert them into each layer of ViT respectively.

[0012] S12 Dataset partitioning: Divide the categories of the training dataset into 10 equal parts, and train one set of categories for each task. For a dataset with N c different categories where represents the number of incremental tasks, N inc is the number of classes to be trained in each incremental task stage, refers to the dataset for each incremental task. After the t-th learning task is completed, all learned categories are represented as where Y(t) represents the set of class labels in each incremental task dataset.

[0013] S13 Initialize semantic prompt: The present invention uses semantic prompts to enhance the generalization ability of visual features. The semantic description R inc for a task's classes is:

[0014] "A photo of {class1} or {class2}…or {classN inc}."

[0015] where class1, class2,..., classN inc will be replaced by the class names in the task class label set Y(t). Then, use the pre-trained large language model BERT as a semantic feature extractor to convert the semantic description R inc into a text embedding:

[0016] P s = BERT(R lan ),

[0017] where represents the task-level semantic prompt, indicating that the dimension of P s is a constant d.

[0018] S14 Initialize visual prompt: The visual prompt is defined as where L vp is the length of the visual prompt, i.e., indicating that the dimension size of the visual prompt P v is L vp × d. The visual prompt P vRandomly initialize at the beginning of training. Both visual cues and semantic cues are task-level cues and will be trained in their corresponding tasks. All trained cues will be stored in the cue pool.

[0019] S15 Initialize keys: In each task, the cues initialize a corresponding set of learnable vector keys for selecting appropriate cues from the cue pool. The number of keys in each set is the same as the number of categories learned in that task, i.e., each category corresponds to one key. The set of trainable keys in the model is denoted as where represents the set of keys in the t-th task.

[0020] Step 2: Model with a multi-mechanism query-key matching method

[0021] S21 Extract query features: Input the query image into the frozen ViT to extract query features It is denoted that the feature dimension of q is d.

[0022] S22 Calculate the score vector: Calculate the cosine similarity between the query feature q and the keys of all categories in the current training task in turn, and combine the results into a score vector

[0023]

[0024] where, represents the N inc -th key vector in the t-th task.

[0025] S23 Calculate the probability distribution: Convert the score vector to a probability distribution l through the softmax function:

[0026]

[0027] S24 Calculate the matching loss: To make the query feature as similar as possible to the key corresponding to the class label, while minimizing the similarity to the keys of other classes, the cross-entropy loss function is used to train the feature distribution of the keys:

[0028]

[0029] where, represents the cross-entropy loss value of the keys, and y i represents the label of the corresponding sample.

[0030] S25 Calculate the entropy of the probability distribution: During inference, when the categories in other tasks are similar to those in the current task, this multi-key matching mechanism may select the cues of other tasks, resulting in bias in image classification. To address this issue, a method introduces an entropy-based auxiliary matching approach. After calculating the probability distribution of the query features in each task, the entropy value H(l) of each distribution is calculated:

[0031]

[0032] S26 Calculate the probability distribution and its entropy of the query prototype: To further improve the accuracy of cue selection, the present invention calculates the average value of the features of each query category as the prototype During inference, the softmax function is used to process the similarity scores between the query features and all prototypes of the t-th task. In this way, a new probability distribution vector ζ can be obtained c :

[0033]

[0034] Similarly, its entropy value H(ζ) is also used to improve query key matching:

[0035]

[0036] S27 Integrate four matching strategies of the module: The model integrates four strategies for cue selection. During inference, the module will execute each strategy in turn:

[0037]

[0038] Among them, P 1 represents the prediction result obtained through the probability distribution l of the t-th task t obtained, P 2 the prediction result obtained through the entropy value H(l t ) of the probability vector of the t-th task, P 3 represents the prediction result obtained through the prototype probability distribution vector ζ of the t-th task t obtained, P 4 represents the prediction result obtained through the entropy value H(ζ t ) of the prototype probability distribution vector of the t-th task.

[0039] S28 Obtain the matching result: The voting strategy is adopted to select the final task cue. That is, the prediction results of these four strategies are counted, and the value that appears most frequently in P 1 , P 2 , P 3 , P 4 is selected. When all the outputs are different, the first output is defaultly selected as the final prediction result.

[0040] S29 Combined Model Input: The prompts of the predicted tasks are taken from the prompt pool to form the input of the ViT backbone network. The input consists of the following multiple parts:

[0041]

[0042] Among them, ξ cls is the class token, and ξ e refers to the image token. P s and P v represent semantic prompts and visual prompts respectively. represents the input information of the ViT backbone network, and its dimension is calculated from the number of image embedding vectors and the constant d.

[0043] Step 3: Set the loss function

[0044] S31 Cosine Contrast Loss: In deep learning, contrastive loss is often used to learn the similarity and difference between features, which can effectively prompt the model to learn that positive sample pairs have a smaller Euclidean distance in the feature space, while keeping negative sample pairs at a larger Euclidean distance. In the method of this paper, in order to ensure that the language prompt can guide the model to learn the category information, the present invention proposes a new cosine contrast loss to make the language prompt output by the network as similar as possible to the category-level language features of the corresponding labels. Specifically, first, the visible category names in the dataset are formed into category-level language expressions R sem,i :

[0045] R sem,i = "A photo of {classi}."

[0046] To ensure the consistency of the language embedding space, Bert is also used for encoding:

[0047] ξ sem,i = Bert(R sem,i )

[0048] Since there are significant differences in categories and domains in the images in the dataset, if the Euclidean distance is directly used to measure the difference between features, the loss value may fluctuate greatly due to the differences in the dataset. To enhance the stability of the model learning process, the present invention proposes to use the cosine contrast loss function. This method can not only output a more stable loss value, but also better adapt to the diversity in the dataset, thereby improving the performance and robustness of the model. The cosine contrast loss proposed by the present invention can be expressed in the following form:

[0049]

[0050] In the incremental task, the number of negative sample pairs is often much larger than that of positive sample pairs. For example, in a 10-class incremental stage, there is only one positive sample pair while there are nine negative sample pairs. The imbalance between the number of positive and negative samples will lead to a greater impact of the loss of negative sample pairs, and it is difficult for the model to learn knowledge from positive samples. To eliminate this impact, a deviation parameter α is set in the function and is set to 0.3 according to experience. In addition, when measuring the similarity between features using cosine similarity, a similarity value between [0,1] indicates similarity, and a value between [-1,0] indicates dissimilarity. When the impact of negative sample pairs is too large, a negative loss value may occur. Optimization algorithms such as gradient descent rely on the gradient of the loss function to update model parameters. If the loss value is negative, the calculated gradient may be incorrect, which will cause the optimization process to deviate from the correct direction. To avoid this situation, the absolute value of the similarity of negative sample pairs is taken to ensure that the calculated cosine contrast loss is always a positive value. By minimizing the contrast loss, the model can learn a feature representation that makes the language prompt processed by the network as similar as possible to its corresponding class-level language feature and as dissimilar as possible to the language features of other classes, thereby promoting the language prompt to better guide the model to learn knowledge and improve the classification performance of the model.

[0051] S32 Classification loss: Each task has its specific trainable classifier Φ(·; φ), which predicts the class of the image based on the feature ξ of the image. cls The present invention uses cross-entropy loss to optimize the classifier:

[0052]

[0053] S33 Total loss function: Generally speaking, for task t, the final loss function combines the multi-key loss cosine contrast loss and CE classification loss to form the overall learning objective of the proposed model:

[0054]

[0055] Step Four: Adjust the network parameters and perform the training and testing of the neural network.

[0056] S41 Parameter Settings: The Pytorch architecture is adopted and the training is carried out on the NVIDIA RTX 3090 GPU. The optimizer uses the Stochastic Gradient Descent (SGD) optimizer with a momentum of 0.9 and an initial learning rate of 0.01. During training, the batch size for each iteration is 24 images. The learning rate decays to zero according to the cosine annealing schedule. The number of training epochs is set to 14 on ImageNetR and 20 epochs on CIFAR100 and ImageNetA.

[0057] S42 Model Performance Testing: Three metrics widely used in continual learning are used to evaluate the model. The average prediction accuracy of all classes after the last training is denoted as Last-acc, the average accuracy of all classes is denoted as Avg-acc, and the forgetting fraction is denoted as FF. To avoid the contingency of experiments and ensure reliable results, the experiments are repeated after shuffling the class order in the dataset using three different random seeds. The final results are the average performance metrics of the three runs and the corresponding standard deviations.

[0058] Compared with the prior art, the present invention has achieved better performance on three publicly available continual learning datasets, significantly outperforming the prior methods in terms of the Last-acc, Avg-acc, and FF metrics. The present invention has not only achieved excellent performance in the common 10 continual learning tasks, but also achieved the best results in the 20 continual learning task settings, fully demonstrating the excellent stability of the present invention in long-sequence tasks. Description of the Drawings

[0059] Figure 1 It is a specific implementation flowchart of the present invention.

[0060] Figure 2 It is a performance comparison of the present invention on three datasets, ImageNetR, CIFAR100, and ImageNetA, under the 10 continual learning task settings. The first four rows are prompt-based methods, the next two rows are adapter-based methods, and the AESP proposed by the present invention is in the last row.

[0061] Figure 3 It is a trend chart of the test results of the present invention and other continual learning methods after each task.

[0062] Figure 4 It is a performance comparison of the present invention on three datasets, ImageNetR, CIFAR100, and ImageNetA, under the 20 continual learning task settings. The first four rows are prompt-based methods, the next two rows are adapter-based methods, and the AESP proposed by the present invention is in the last row.

[0063] Figure 5For the ablation experiment of the AESP method proposed by the present invention, the effectiveness of the adapter, semantic prompt, and fusion multi-mechanism query-key matching method was verified respectively. Detailed implementation mode

[0064] The present invention conducts detailed experimental verification on the above methods and evaluates the effect of model continuous learning on three datasets: ImageNetR, CIFAR100, and ImageNetA. ImageNetR consists of 30,000 images from 200 categories of the ImageNet dataset. Each category includes images of various styles, such as art, cartoons, and graffiti, thus introducing significant image style variations. This style diversity makes the dataset challenging. CIFAR-100 contains 60,000 color images with a resolution of 32×32 pixels, divided into 100 categories, with 500 training images and 100 test images in each category. It is a recognized benchmark dataset in the field of continuous learning. ImageNetA is a real-world dataset containing 7,500 unmodified naturally occurring images from 200 ImageNet categories. These images are all misclassified by the ResNet model, posing a great challenge to machine learning models.

[0065] To evaluate our method in different continuous learning settings, we adopted two strategies for ImageNetR and ImageNetA, each strategy containing 200 categories. The first strategy divides the dataset into 10 tasks, each task containing 20 categories; the second strategy divides the dataset into 20 tasks, each task containing 10 categories. For CIFAR100 containing 100 categories, we divide it into 10 tasks, each task containing 10 categories.

[0066] Figure 2Shows the evaluation results of the present invention on three datasets. It can be clearly seen from the results that the results of the present invention far outperform the prompt-based methods. Although Cprompt has made good progress compared to the previous prompt-based methods, our method still outperforms it by 5.02% and 3.61% respectively in terms of Last-acc and Avg-acc on imagenetR, and in terms of the forgetting rate, there is a 0.91% relief compared to it. In addition, although the latest adapter-based methods have once surpassed the prompt-based methods in terms of results, the AESP proposed by the present invention has surpassed them again. Specifically, it outperforms the best adapter-based method - SSIAT by 2.05% and 2% respectively in terms of Last-acc and Avg-acc. Different from the result that the present invention leads more in Last-acc among the prompt-based methods, the leading advantage of the present invention in Avg-acc is greater than that in Last-acc, because the adapter-based methods have better anti-forgetting ability. And the method of the present invention that combines prompt and adapter well inherits the advantages of the adapter, making the forgetting rate decrease by about 1.38% compared to SSIAT.

[0067] Due to the particularity of the imagenetA dataset, the results on this dataset are generally relatively low. Nevertheless, the method of the present invention has achieved 63.99% and 72.01% respectively in terms of Last-acc and Avg-acc, which is about 1.56% and 1.18% higher than SSIAT.

[0068] Figure 3 Shows the trend charts of the image classification accuracy in each task for the present invention and the comparative methods. These curves are plotted based on the average results of three seeds in each consecutive learning stage. The curve of the present invention is always above all other curves, reflecting the excellent recognition accuracy and anti-forgetting ability of the present invention.

[0069] Figure 4 Shows the comparative experiment results under 20 consecutive learning task settings. In contrast, the present invention has achieved the best results, which fully demonstrates that the present invention also has excellent stability in long-sequence tasks.

[0070] Figure 5Shows the ablation experiment conducted by the present invention, aiming to evaluate the effectiveness of three key components: the adapter, semantic cues, and the fusion of multi-mechanism query-key matching method. The experimental results show that removing any one of the components will lead to a decline in model performance, which is in line with the expected experimental results and fully demonstrates the importance and effectiveness of the components of the present invention.

Claims

1. A continuous learning method based on adapter-enhanced semantic cues, characterized in that: The implementation steps of this method are as follows: Step 1: Model initialization; S11 Model loading: load the pre-trained ViT and large language model BERT weights into the corresponding network structure and freeze them; initialize the adapter and insert it into each layer of ViT respectively; S12 divides the data set: the categories of the training data set are evenly divided into 10 parts, and each task trains a group of categories; S13 Initialize semantic hints: Use semantic hints to enhance the generalization ability of visual features; semantic description of a task category R inc for: ″A photo of{class1}or{class2}…or{classN inc }.″ Among them, class1, class2, ..., classN inc Replace it with the class name in the task category label set Y(t); use the pre-trained large language model BERT as the semantic feature extractor to extract the semantic description R inc Convert to text embedding: P s =BERT(R lan ), in, It represents the task-level semantic hint. Indicates P s The dimension is a constant d; S14 Initialize visual cues: The visual cues are defined as Where L vp is the length of the visual cue, That is, the visual cue P v The dimension size is L vp ×d; visual cue P v Initialized randomly at the beginning of training; S15 Initialize key: Each prompt in the task initializes a set of corresponding learnable vector keys to select appropriate prompts from the prompt pool. The number of keys in each set is consistent with the number of categories learned in the task, that is, each category corresponds to a key; the set of trainable keys in the model is expressed as in Represents the set of keys in the tth task; Step 2: Modeling by integrating multi-mechanism query-key matching methods S21 Extract query features: Input the query image into the frozen ViT and extract query features Indicates that the feature dimension of q is d; S22 calculates the score vector: calculates the cosine similarity of the query feature q and the key of all categories in the current training task in turn, and combines the results into a score vector in, Indicates the Nth task in the tth task inc key vector; S23 calculates the probability distribution: the score vector is converted into a probability distribution l through the softmax function: S24 calculates the matching loss: In order to make the query feature similar to the key corresponding to the category label, while minimizing the similarity with the key of other categories, the cross entropy loss function is used to train the feature distribution of the key: in, Represents the cross entropy loss value of key, y i Indicates the label of the corresponding sample; S25 calculates the entropy of probability distribution: During reasoning, an entropy-based auxiliary matching method is introduced; after calculating the probability distribution of the query feature in each task, the entropy value H(l) of each distribution is calculated: S26 calculates the probability distribution and entropy of the query prototype: To improve the accuracy of the selection prompt, the feature average of each query category is calculated as the prototype During inference, the softmax function is used to process the similarity scores between the query feature and all prototypes of the t-th task to obtain a new probability distribution vector ζ c : Similarly, its entropy value H(ζ) is also used to improve query key matching: S27 integrates four matching strategies: Integrates four strategies for selecting prompts. During reasoning, the module will execute each strategy in turn: Among them, P1 represents the probability distribution l of passing the tth task t The prediction result obtained, P2, is obtained by the probability vector entropy value H(l t ) obtained, P3 represents the prototype probability distribution vector ζ of the t-th task t The prediction result obtained, P4 represents the entropy value H(ζ) of the prototype probability distribution vector of the t-th task t ) The prediction results obtained; S28 obtains the matching result: adopts a voting strategy to select the final task prompt; that is, the prediction results of the four strategies are counted, and the value that appears most frequently in P1, P2, P3, and P4 is selected. When all the outputs are different, the first output is selected as the final prediction result by default; S29 Combined model input: The predicted task prompts are taken from the prompt pool to form the input of the ViT backbone network. The input consists of the following parts: Among them, ξ cls is the class label, ξ e Refers to the image tag, P s and P v Represent semantic cues and visual cues, respectively; Represents the input information of the ViT backbone network, its dimension The number of image embedding vectors Calculated from the constant d; Step 3: Set the loss function S31 Cosine contrast loss: A new cosine contrast loss is proposed to make the language prompts output by the network as similar as possible to the category-level language features of the corresponding labels; the category names visible in the dataset are combined into category-level language expressions R sem,i : R sem,i ="A photo of{classi}.″ In order to ensure the consistency of the language embedding space, Bert is also used for encoding: x sem,i =Bert(R sem,i ) The cosine contrast loss function is used to improve the performance and robustness of the model; the proposed cosine contrast loss function It is expressed as: S32 classification loss: Each task has its own specific trainable classifier Φ(·;φ), which is trained according to the image features ξ cls To predict the category of the image; using cross entropy loss To optimize the classifier: S33 Total loss function: In general, for task t, the final loss function combines the multi-key loss Cosine contrast loss and CE classification loss The overall learning objectives of the proposed model are formed: Step 4: Adjust network parameters, train and test the neural network S41 parameter settings: Pytorch architecture is used, and the optimizer uses the stochastic gradient descent SGD optimizer with a momentum of 0.9 and an initial learning rate of 0.01; during training, the batch size of each iteration is 24 images; the learning rate decays to zero according to the cosine annealing schedule; the iteration cycle for training on ImageNetR is set to 14, and the iteration cycle for training on CIFAR100 and ImageNetA is set to 20; S42 model performance test: Three indicators widely used in continuous learning are used to evaluate the model; the average prediction accuracy of all categories after the last training is recorded as Last-acc, the average accuracy of all categories is recorded as Avg-acc, and the forgetting score is recorded as FF.

2. A continuous learning method based on adapter-enhanced semantic cues according to claim 1, characterized in that: In S12, for N c Datasets of different categories in Indicates the number of incremental tasks, N inc is the number of classes to be trained at each incremental task stage, refers to the data set of each incremental task; after the tth learning task, all learned categories are represented as Among them, Y(t) represents the category label set in each incremental task dataset.

3. The continuous learning method based on adapter-enhanced semantic cues according to claim 1, characterized in that: In S14, both visual cues and semantic cues are task-level cues, which will be trained in their corresponding tasks, and all trained cues will be stored in a cue pool.

4. The continuous learning method based on adapter-enhanced semantic cues according to claim 1, characterized in that: In S42, in order to avoid the contingency of the experiment and ensure the reliability of the results, three different random seeds are used to shuffle the order of categories in the data set and then repeat the experiment; the final result is the average performance index of the three runs and the corresponding standard deviation.

Citation Information

Cited By

  • Visual language bimodal model end cloud collaborative inference system based on low-rank fine tuning

    CN122044859A