A Cross-Subspace Knowledge Alignment Method for Continuous Learning of Image Classification

By constructing a network model that is aligned across subspace knowledge in class incremental learning, using the alignment loss function of features and decision boundaries, the feature subspace misalignment caused by task recognition error is solved, and the model is robust and plasticized to incorrect task identification.

CN119942270BActive Publication Date: 2025-06-17XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510421323.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-17
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the category incremental learning scenario, due to the recognition error of the task recognizer, the test image may be projected into the wrong feature subspace, resulting in blurred classification results and lack of robustness.

Method used

A cross-subspace knowledge alignment method is proposed. By constructing a network model, including feature extractor, task sharing prompt vector, independent task submodule and classifier between tasks, the alignment loss function at the feature level and decision boundary level is used to realize the alignment and aggregation of cross-subspace knowledge.

Benefits of technology

It significantly improves the robustness of the model for wrong task identification, can ensure the reliability and accuracy of the model when task identification is misleading, and maintains the plasticity of the model, adapting to the complex needs of multi-task learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942270B_ABST
    Figure CN119942270B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-subspace knowledge alignment method for continuous learning image classification, aiming to enhance the robustness of the model against misleading task IDs. The cross-subspace knowledge alignment and aggregation method of the present invention, while learning the task recognizer, also ensures the plasticity of the model through a dedicated task sub-module and adopts a dual knowledge alignment training framework. The dual knowledge alignment aligns the feature semantics and decision boundaries of different sub-spaces at the feature and classifier decision boundary levels, thereby overcoming the ambiguous decision problem caused by sub-space misalignment. In addition, in order to avoid overconfidence in misleading task IDs during the inference process, the present invention proposes a powerful inference mechanism of task confidence-guided adapter mixing, which achieves more robust inference through the soft knowledge aggregation of the task sub-module. Experiments show that the cross-subspace knowledge alignment and aggregation method of the present invention outperforms the existing class incremental learning methods based on efficient parameter fine-tuning in terms of performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of continuous learning image classification, and particularly relates to a cross-subspace knowledge alignment method for continuous learning image classification. Background Art

[0002] Continuous learning requires the model to continuously learn from continuously arriving data streams to adapt to the ever-changing real-world scenarios. This is crucial for the deployment of artificial intelligence models in practical applications. However, continuous learning faces a major challenge: during the process of learning new classes, the model often encounters the catastrophic forgetting problem, that is, the knowledge of old classes gradually gets lost. Currently, there are mainly two common scenarios in continuous learning: task-incremental learning (TIL) and class-incremental learning (CIL). In the TIL scenario, the task identifiers (task-ids) can be accessed during the inference phase, and this setting is considered too simplistic. The CIL scenario is more challenging because the task identifiers are unknown during inference, and the model needs to make inferences without explicit task identifiers.

[0003] With the development of pre-trained models, methods based on parameter-efficient fine-tuning (PEFT) have shown excellent performance in visual tasks. These methods train by inserting lightweight modules on the basis of freezing the pre-trained backbone network, and can reduce the computational cost and training overhead while maintaining the generalization ability of the pre-trained model. In recent years, PEFT methods have been introduced into the field of continuous learning, especially showing strong performance in the class-incremental learning scenario. These methods capture task-specific knowledge by assigning independent sub-modules to each task, and provide sufficient model plasticity for new tasks. Existing PEFT methods usually design task recognizers to identify task identifiers in class-incremental learning, and select the corresponding sub-modules to classify test images. However, due to the inevitable recognition errors of the task recognizers, the test images may be projected into the wrong sub-module subspace. In this case, since the existing methods independently train the sub-modules and their classifiers, they fail to establish effective interactions between the feature subspaces of different sub-modules, resulting in the problem of feature sub-space misalignment. This problem is mainly manifested in two aspects: (1) there is an inconsistency in the semantics of the same image features from the correct sub-module and the wrong sub-module; (2) the decision boundary trained based on the correct task identifier cannot effectively classify the features in the wrong subspace.

[0004] Based on the above analysis, the main issues to be considered for achieving this task are as follows: Due to the recognition errors of the task recognizer, the test images may be projected into the wrong feature subspaces, resulting in ambiguous classification results. Therefore, the core problem to be solved is how to effectively align and aggregate the feature subspaces of different sub-modules to improve the robustness of the model in the case of task recognition errors. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention provides a cross-subspace knowledge alignment method for continuous learning image classification. The technical problems to be solved by the present invention are achieved through the following technical solutions:

[0006] A cross-subspace knowledge alignment method for continuous learning image classification includes:

[0007] S1. Obtain a continuous learning training data set for the class incremental learning scenario. The continuous learning training data set includes image data sets for multiple tasks of image classification, and the image data sets of all tasks have different and non-overlapping categories;

[0008] S2. Build a network model for continuous learning image classification. The network model includes: a feature extractor formed by freezing the parameters of the backbone network of a pre-trained ViT network, L trainable task-shared prompt vectors, trainable task-independent task sub-modules, and a classifier constructed by a fully connected layer ; wherein, the feature extractor contains L transformer blocks, and each prompt vector is inserted into the original input of a transformer block so that the L prompt vectors and the feature extractor together form a task-sharing network; each task sub-module includes L adapters, which act on the L transformer blocks of the feature extractor respectively;

[0009] S3. Determine the basic loss function of the network model, and construct a loss function for feature-level alignment and a loss function for decision boundary-level alignment. Use the sum of all loss functions to construct a total loss function, thereby constructing a training framework for cross-subspace knowledge alignment;

[0010] S4. Under the training framework, use the image data sets of the multiple tasks to train the network model in sequence until the value of the total loss function converges, and obtain a trained network model;

[0011] S5. Input the test data set into the trained network model, and realize cross-subspace knowledge aggregation through multi-adapter hybrid inference guided by task confidence, thereby realizing the continuous learning image classification task.

[0012] Advantages of the present invention:

[0013] 1. The present invention provides a method for cross-subspace knowledge alignment and aggregation, which significantly enhances the robustness of the model to incorrect task identifiers (task-ids). This feature is crucial for class-incremental learning (CIL) methods based on parameter-efficient fine-tuning (PEFT), as it can ensure the reliability and accuracy of the model even when the task identifiers are misleading;

[0014] 2. The present invention is particularly suitable for handling long-sequence tasks and can maintain the plasticity of the model even as the number of tasks increases. This means that regardless of the number of consecutive tasks, the model can flexibly adapt to new tasks while effectively retaining knowledge of the learned tasks, thereby achieving better performance in the long-term learning process.

[0015] Through this design, the method of the present invention can not only address the challenge of incorrect task identifiers in the short term but also adapt to the complex requirements of multi-task learning in the long term, providing an efficient, robust, and flexible solution for continuous learning scenarios in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic flow chart of a cross-subspace knowledge alignment method for continuous learning image classification provided by an embodiment of the present invention;

[0017] Figure 2 is a schematic diagram of the framework and principle of the network model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0019] To effectively align and aggregate the feature subspaces of different sub-modules and enhance the robustness of the model in the case of incorrect task recognition, an embodiment of the present invention provides a cross-subspace knowledge alignment method for continuous learning image classification, as Figure 1 shown, the method may include the following steps:

[0020] S1. Obtain a continuous learning training data set for the class-incremental learning scenario, where the continuous learning training data set includes image data sets for multiple tasks for image classification, and the image data sets of all tasks have different and non-overlapping categories;

[0021] In an embodiment of the present invention, the continuous learning training data set is represented as:

[0022] ;

[0023] wherein, the continuous learning training data set includes The image dataset of the task, is a natural number greater than 0; the image dataset of the task is denoted as , and the size of is ; denotes the th image data in ; ; denotes 's class label; denotes the label space of task , and the label spaces of different tasks are disjoint; , denote the height and width of the image data, denotes the number of channels of the image data, , and are all natural numbers greater than 0.

[0024] In the embodiments of the present invention, all tasks are image classification tasks, but the number of classification categories can be different. For example, it can be a binary classification task or a multi-classification task.

[0025] S2. Build a network model for continuous learning of image classification. The network model includes: a feature extractor formed by freezing the parameters of the backbone network of a pre-trained ViT network, L trainable task-shared prompt vectors, a trainable task-independent task sub-module, and a classifier constructed by a fully connected layer ; wherein, the feature extractor contains L transformer blocks, and each prompt vector is inserted into the original input of a transformer block so that the L prompt vectors and the feature extractor together form a task-sharing network; each task sub-module includes L adapters, which act on the L transformer blocks of the feature extractor respectively;

[0026] Please refer to Figure 2 for the framework and principle schematic diagram of the network model shown. In the embodiments of the present invention, the backbone network of Vision Transformer (ViT) is pre-trained based on the ImageNet-21k dataset, and after the training is completed, the parameters are frozen to obtain the feature extractor; the feature extractor contains L transformer blocks, and each transformer block is composed of a MSA (Multi-head self-attention) layer and an FFN (Feed-Forward Neural Network) with residual connections. For simplicity, Figure 2The complete structure of the L transformer blocks is not shown, and only the MSA and FFN are given to illustrate the main components of the transformer block. Figure 2 In this context, the task-sharing prompt refers to the prompt vector shared among trainable tasks, which is achieved based on visual prompt tuning; the task-independent adapter refers to the task sub-module independent among trainable tasks, Each of the tasks corresponds to a task sub-module; the classifier refers to the classifier constructed by the fully connected layer .

[0027] To facilitate the understanding of the structure and processing principle of the network model in the embodiments of the present invention, the feature extractor will be described first.

[0028] Specifically, before inserting the L prompt vectors, the feature extractor is used to:

[0029] (1) Divide the input image data into patch blocks, and encode each patch block into a -dimensional vector to obtain the block-encoded features of the input image data ;

[0030] Among them, the input image data , , represent the height and width of the image data, represents the number of channels of the image data; and are natural numbers greater than 0, represents the dimension of the image features;

[0031] (2) Add the classification token and the position embedding vector to the block-encoded features to obtain the image patch-encoded representation of the input image data ,

[0032] where, , , (1);

[0033] The above two steps can be understood as the pre-processing process of the feature extractor before the L transformer blocks, which can be implemented by a patch embedding layer (Patch layer), which is a learnable linear projection layer that will evenly divide the input image data (i.e., the input image) into blocks, and each block is named a patch block here. For Each of the patch blocks is encoded and mapped into a vector of dimensions, thus transforming the input image data into block-encoded features

[0034] The classification token is a learnable classification vector with initialization information when set, containing multiple elements corresponding to the confidence levels of each category. Through the model learning process, the output of the feature extractor will ultimately be formed.

[0035] The position embedding vector represents the relative position information of all elements in

[0036] The processing process of the Patch layer can be understood with reference to related technologies and will not be described in more detail here.

[0037] (3) Process the encoded representation of the image patches through L transformer blocks of the feature extractor, where each transformer block consists of a multi-head self-attention (MSA) layer and a feed-forward network (FFN) with residual connections;

[0038] For the input to the th transformer block, the forward propagation process of this transformer block is expressed by the formula:

[0039] (2);

[0040] where, represents the output of the MSA layer in the th transformer block, serving as the input to the FFN in the th transformer block, represents the input to the th transformer block; and represent the processing of the MSA layer and the FFN respectively.

[0041] That is to say, for the th transformer block, its input is processed by its own MSA layer, outputting , which enters its own FFN for processing, outputting as the input to the th transformer block.

[0042] As can be understood by referring to the previous text, the input to the first transformer block is the encoded representation of the image patches , that is at the time of , then the output of the first transformer block, which is the input of the second transformer block, is , and so on. The output of the L-th transformer block is .

[0043] (4) Use the classification token marking features output after L transformer blocks as the input image data The image features output by the feature extractor are expressed as:

[0044] (3);

[0045] Among them, represents the feature extractor; represents extracting the element at the first position in the output after L transformer blocks , that is, extracting the at this time.

[0046] The embodiment of the present invention develops a task-sharing network based on Visual Prompt Tuning (VPT). Since visual prompts can inject knowledge as a global operation before each transformer block, they are very suitable for capturing the common knowledge of downstream tasks and have strong generalization ability.

[0047] Next, the L prompt vectors and the task-sharing network formed by inserting the L prompt vectors after the feature extractor and the feature extractor will be described.

[0048] Specifically, the L prompt vectors are a set of task-sharing learnable vectors, and their set is represented as . After inserting the L prompt vectors, the task-sharing network is used for:

[0049] For the -th transformer block, the inserted prompt vector is concatenated with the input of this transformer block as the input of the MSA layer of this transformer block, thereby changing the forward propagation process, and outputting shared features via the last layer of the task-sharing network;

[0050] Among them, the forward propagation process of the task-sharing network is expressed by the formula:

[0051] ; (4);

[0052] Among them, the set of L prompt vectors is represented as , where the -th prompt vector Inserted into the original input of the th transformer block, , indicating the length of the inserted prompt vector, indicating the dimension of the image features, being a natural number greater than 0; indicating the output features of the MSA layer of the th layer of the task sharing network; indicating the input of the th layer of the task sharing network;

[0053] The shared features output by the task sharing network are expressed by the formula:

[0054] (5);

[0055] Compared with the separate feature extractor in formula (2), the input of the MSA layer in formula (4) is based on the original input and adds the inserted prompt vector , making the output of the MSA layer become , and the in this output is still input into the corresponding FFN for processing. It can be understood that from formula (2) to formula (4), although both represent the output of the th transformer block, there are actually changes.

[0056] Combining formulas (3) and (5), it can be seen that before inserting L prompt vectors, the output of the feature extractor is , and after inserting L prompt vectors, the output of the feature extractor (which actually becomes the task sharing network at this time) is .

[0057] Since the task sharing network needs to adapt to all downstream tasks, when training a new task, it is necessary to reduce its forgetting phenomenon. To achieve this goal, the present invention updates along the direction orthogonal to the previous feature space . will be updated in all consecutive tasks, encapsulating the common knowledge between them. Through this orthogonally optimized prompt, a shared backbone network and a powerful task recognizer are established, specifically tailored for downstream tasks.

[0058] Due to the strict orthogonal constraint, the task sharing network is not sufficient to capture the fine-grained knowledge specific to each individual task, thus limiting the plasticity of the model. Therefore, the present invention assigns a task-specific task sub-module to each task to capture its unique knowledge and generate more discriminative features. The task sub-module of the present invention is designed based on the adapter architecture.

[0059] For a specific task the task task sub-module, denoted as ;

[0060] Among them, is the th adapter in, acting on the th layer of the feature extractor. Due to the residual connection property of adapter tuning, it can effectively enable the model to independently aggregate unique knowledge from different tasks.

[0061] Consists of a downsampling matrix and an upsampling matrix ; represents the intermediate feature dimension of the adapter, ;

[0062] The forward propagation process of

[0063] (6);

[0064] represents the input of represents the activation function;

[0065] In the embodiments of the present invention, each task sub-module includes L adapters, corresponding to acting on L transformer blocks of the feature extractor, in the following manner:

[0066] The task the th adapter in the task sub-module of , integrates the knowledge corresponding to the task into the th FFN of the task sharing network through a residual connection. The corresponding forward propagation process is expressed as:

[0067] (7);

[0068] Referring to in formula (4), after adding the task sub-module, it becomes formula (7), with the addition of part.

[0069] S3, determining the basic loss function of the network model, and constructing a feature-level aligned loss function and a decision boundary-level aligned loss function, and constructing a total loss function using the sum of all loss functions, thereby constructing a cross-subspace knowledge alignment training framework;

[0070] First, determine the basic loss function of the network model:

[0071] In the During the task, a training dataset of size Batch (Superscripts omitted for simplicity ), from the shared network and Extract features from task submodules and ,in , The output features of the network are then passed through a classifier constructed by a fully connected layer. ,for For classification problems, the dimension of the fully connected layer is . Network features and classifiers The commonly used cross entropy loss is used for training.

[0072] The basic loss function of the network model is expressed as:

[0073] (8);

[0074] in, Represents the basic loss function of the network model. Figure 2 In the expression ; Indicates the current task Batches of image datasets size; Represents the softmax activation function commonly used in image classification; represents the cross entropy loss; the shared features extracted from the task sharing network are recorded as , will be added from the current task The features extracted by the task submodule are recorded as , Middle Shared Features , Middle Features ; express Through the classifier The characteristics of the latter; express Label; Indicates Features after passing through the classifier ; Indicates Label.

[0075] Given the design of the task sub - modules specific to the tasks of the present invention, it is crucial to estimate the task ID through the task recognizer to select the appropriate sub - module. However, due to the limited performance of the task recognizer, there are inevitably incorrect task IDs. To enhance the robustness of the model to incorrect task IDs, a two - layer knowledge alignment (DKA) training method is proposed, which aligns task - specific knowledge at the feature and classifier levels.

[0076] (I) Feature - level alignment:

[0077] Generally speaking, during inference, it is only necessary to combine the feature extractor, the prompt vector, and the task sub - module corresponding to the test image to generate image features. However, since the task label is unknown during inference, it is very likely to assign the test image to the task sub - module corresponding to the wrong task label. And different task sub - modules will project the same image into different feature sub - spaces.

[0078] For a robust feature extractor that can generate meaningful features, even if the task ID is incorrect, the same image should achieve consistent semantics when projected into the correct and incorrect feature sub - spaces. In this regard, the present invention unifies the semantics of features in different feature sub - spaces through feature - level alignment. Specifically, when the same class of images pass through the networks constructed by different task sub - modules, they should also output as similar features as possible.

[0079] The present invention proposes a cross - subspace feature alignment (CSFA) loss, which encourages the features of the same class in the current subspace to align with the features in the previous subspace, that is, the samples of the same class within the current task will output similar features after passing through the task sub - module of the current task and the task sub - module of the past task.

[0080] Because the training process of the network model is carried out for multiple tasks in sequence, that is, successively using for corresponding task training; first, for the images in the current image dataset , extract the features of these images in the th subspace ([[]] [[]]), that is, for the current task, pass its image dataset through the task sub - module of the current task and the task sub - module of the past task respectively to obtain the image features, denoted as and . During training, a feature set is randomly selected from the previous subspace. ​​

[0081] Given an input batch , for the features with label in the current subspace, construct its positive set through the features from the sampling set with the same ground truth, where . Meanwhile, regard the features with different labels in the current subspace as negative samples, where .

[0082] That is to say, for images belonging to the same class, the present invention regards their features in the current subspace and the past subspace as positive sample pairs, and narrows the distance between them in the loss function. On the contrary, for images belonging to different classes, their features in the current subspace are regarded as negative sample pairs, and the distance between them is widened in the loss function.

[0083] For the input batch , its CSFA loss function, that is, the loss function for feature-level alignment, is expressed as:

[0084] (9);

[0085] where represents the loss function for feature-level alignment, which is used to measure the cross-subspace feature alignment loss of the input batch , by narrowing the distance between the features of the same-class images in different feature subspaces and widening the distance between the features of different-class images in the same feature subspace, the features of the same-class samples in multiple subspaces are aggregated, so as to achieve the semantic alignment of the same-class features; represents the positive samples in different task feature spaces; represents the positive sample set; represents the negative sample set; represents all the features in a batch except ; represents the temperature factor for contrastive learning; represents taking the union; represents the natural exponential function; represents the cosine similarity function.

[0086] The CSFA loss narrows the distance between the features of the same-class images in different feature subspaces and widens the distance between different images in the same feature subspace, and aggregates the features of the same-class samples in multiple subspaces, thus achieving the semantic alignment of the same-class features.

[0087] (II) Decision boundary level alignment:

[0088] Due to the inevitable subspace gap, the classifier trained only with the correct task ID cannot handle features within the incorrect subspace. Decision boundary alignment aims to learn a clear decision boundary for these misprojected features. Specifically, in the th task, the present invention trains a task-adaptive classifier which, when the previous data is projected into the current feature subspace, attempts to avoid ambiguous predictions. That is, each task trains its own task-adaptive classifier, which is different from the classifier constructed by the fully connected layer described above . The task-adaptive classifier takes into account the image features of the current task and past tasks, such that the output prediction value of the task-adaptive classifier is not biased towards the current task.

[0089] To achieve this goal, the present invention first simulates the features of the previous data in the current subspace and then jointly trains with the features of the current data using the simulated features . The present invention further decomposes the simulation process into three steps:

[0090] 1. Model the affinity between the current data and the previous data in the shared feature space through a task-sharing network;

[0091] 2. Model the subspace offset in the current data between the current subspace and the shared subspace;

[0092] 3. Transfer the subspace offset of the current shared features to the previous features through the affinity matrix.

[0093] Since modeling the affinity requires the previous data distribution but no samples are saved for training, the present invention solves this problem through Gaussian sampling. After the th task, for the th class in the th task, approximate its feature distribution in the shared feature space as a Gaussian distribution . During the th task, for each arriving batch, sample pseudo-representations from the previous shared feature distribution with a size of , where represents the th pseudo-feature in the sampled set of pseudo-representations, and represents the class label. Combine the current feature and , where represents the a shared feature indicating the class label indicating the th in the features extracted by the task sub-module of the current task the class label indicating the affinity matrix between the previous data and the current data in the shared feature space can be expressed by the following formula:

[0094] (10);

[0095] where represents the affinity matrix the th element in the th row and represents the temperature coefficient; represents the pseudo-representation set the features closest to

[0096] The subspace offset between the current data in the current subspace and the shared feature space can be simply modeled by the residual between a specific feature and the shared feature of the current data . Subsequently, the subspace offset is transferred from the current feature to the sampled pseudo-feature . The simulated feature is expressed as follows:

[0097] (11);

[0098] where represents the simulated feature.

[0099] Then a unified feature set is obtained, which includes the image features of all visible tasks in the current subspace. It is used to train the task adaptive classifier , and the loss function with aligned decision boundary levels is expressed as:

[0100] (12);

[0101] where represents the loss function with aligned decision boundary levels for training the task adaptive classifier ; represents the unified feature set, including the image features of all visible tasks in the current subspace; represents the An adaptive classifier for each task; denote the features in; denote the features after passing through the adaptive classifier ; denote the label of.

[0102] Therefore, for each task, the corresponding adaptive classifier is trained separately. After all the training is completed, in order to reduce the computational overhead of inference, after obtaining the trained network model, the method further includes:

[0103] By aggregating the parameters of the adaptive classifiers of all tasks, a unified classifier for a feature subspace suitable for all tasks is constructed .

[0104] The present invention aggregates the parameters of all task adaptive classifiers, that is , thereby constructing a unified classifier ;

[0105] Aggregating the parameters of the adaptive classifiers of all tasks is implemented using a parameter aggregation formula. For the parameters corresponding to the th category in the th task, that is

[0106] (13);

[0107] where is the weight of the th category in the unified classifier is the bias of the th category in the unified classifier is the weight of the th category in the task adaptive classifier is the bias of the th category in the task adaptive classifier denotes the task number variable, denotes the number of the current task. This aggregation strategy enables the unified classifier to distinguish the features projected into each subspace, even in the face of misleading task IDs.

[0108] Sum the basic loss function, the feature-level alignment loss function, and the decision boundary-level alignment loss function to obtain the total loss function, denoted as:

[0109] (14);

[0110] The training framework for the entire cross-subspace knowledge alignment is trained by minimizing the total loss function.

[0111] For simplicity, the trade-off parameters of all loss terms are set to 1.0. In each training session, the training framework of the present invention can be optimized in an end-to-end manner. This way can keep the decision boundaries of the feature spaces corresponding to different task sub-modules consistent, so that even when the test image is in the situation of unknown task label and wrong selection of the task sub-module, a reliable classification result can still be obtained.

[0112] In summary, the present invention constructs a "cross-subspace knowledge alignment" training framework, which performs double alignment of features and decision boundaries respectively. At the feature level, the features of the same category are aligned in different feature sub-spaces, so that the distributions of the same-category features in cross-task sub-modules tend to be consistent; at the classifier level, a task-adaptive classifier is trained separately for each task, which can classify the features of all learned categories, even if these features are wrongly projected into the feature sub-space of the current task.

[0113] S4. Under the training framework, use the image datasets of the multiple tasks to train the network model in sequence until the value of the total loss function converges, and obtain the trained network model.

[0114] Train the network model according to the training framework and the total loss function therein in S2 and S3, continuously update its parameters, and realize the generation of the optimal network model, that is, obtain the trained network model.

[0115] Specifically, it can be understood by referring to the above training framework and combining the training process of the general continuous learning image classification model, and no detailed description will be given here.

[0116] S5. Input the test dataset into the trained network model, and realize the knowledge aggregation across sub-spaces through multi-adapter hybrid inference guided by task confidence, so as to realize the continuous learning image classification task.

[0117] S5 may include the following steps:

[0118] S51. For any test image data in the test dataset, use the trained network model to obtain the output of the classifier and the shared features, where the output of the classifier contains confidences of classes; is the total number of classifications, and is a natural number greater than 0;

[0119] The test dataset may include test image datasets for multiple tasks.

[0120] In the experiments of the embodiments of the present invention, for the sake of easy understanding, the test dataset is represented as , including test image datasets for tasks; the test image dataset for the th task is represented as , the size of is , represents the th test image data in ; , represents 's class label, represents the label space of task , and the label spaces of different tasks are non - overlapping. , represent the height and width of the image data, represents the number of channels of the image data.

[0121] Input the test dataset into the trained network model for testing and class prediction of images. Given the task - specific nature of the architecture of the present invention, it is crucial to project the test images (i.e., test image data) into an aligned feature space by aggregating task - specific knowledge during the inference process without the task ID. To achieve this goal, the present invention utilizes the task - shared knowledge from the task identifier to aggregate the task - specific knowledge associated with each test image.

[0122] Given a test image data and its shared feature , represent the output of the classifier before softmax as , and represent 's th element as . is the total number of visible classes.

[0123] S52, retain the largest elements in the output of the classifier , and set the remaining elements to infinity to obtain the updated classifier output;

[0124] To extract relevant information from each task, the present invention only retains the classes that are semantically related to the test image and discards other classes. Therefore, only retain 's largest elements that correspond to the categories that are semantically most similar to the test image:

[0125] Specifically, the in elements can be sorted in descending order, and the top elements are selected. Then, the among these elements are retained, and the remaining elements are set to infinity to obtain an updated classifier output.

[0126] The above process can be expressed as:

[0127] (15);

[0128] where represents the updated result of ; represents the top- largest elements in .

[0129] S53, based on the output of the updated classifier , the confidence scores of the shared feature belonging to each task are obtained;

[0130] Specifically, the confidence score belonging to task is obtained, where represents the semantic information amount related to the input contained in each task: (16);

[0131] (16);

[0132] where represents the probability that belongs to task ; is the temperature that controls the smoothness of the task confidence scores.

[0133] S54, for each layer of the task-sharing network, the confidence scores of the shared feature belonging to each task are used as routing weights to aggregate the adapter output features corresponding to the tasks of that layer, and the adapter output features aggregated by all tasks of that layer are integrated as knowledge into the FFN of that layer. After passing through all layers, the aggregated features are output;

[0134] The confidence scores are used as routing weights to aggregate task-specific intermediate features to form a Mixture of Adapters (MoA) architecture. Given the input of the network model as , the The input of a transformer block is , and the forward propagation process is as follows:

[0135] (17);

[0136] Compared with formula (7), formula (17) changes the separate adapter processing result of this transformer block into the weighted result of the adapter result of this transformer block and the corresponding confidence score on all tasks.

[0137] Compared with assigning a single task to each test image, this Mixture of Adapters (MoA) structure avoids overconfident aggregation and further enhances the model's robustness to misleading task IDs.

[0138] The aggregated features after the L-th block are used for final classification, denoted as .

[0139] S55, using the unified classifier and the classifier , determines the predicted class of the test image data.

[0140] Final prediction is jointly determined by the unified classifier and the classifier :

[0141] (18);

[0142] where represents the predicted class;

[0143] The present invention aims to enhance the model's robustness to misleading task IDs, including a knowledge alignment training strategy and a knowledge aggregation inference scheme. By double alignment of feature semantics and decision boundaries and task-adaptive classifier aggregation, it alleviates the subspace misalignment problem caused by task recognition errors and achieves more robust class-incremental learning (CIL). Specifically, cross-subspace knowledge alignment and aggregation, while learning task recognizers, also ensure the plasticity of the model through dedicated task-specific sub-modules and adopt a dual knowledge alignment training framework. Dual knowledge alignment aligns the feature semantics and decision boundaries of different subspaces at the feature and classifier decision boundary levels, thus overcoming the ambiguous decision problem caused by subspace misalignment. In addition, in order to avoid overconfidence in misleading task IDs during the inference process, a task confidence-guided adapter mixture strong inference mechanism is proposed, which achieves more robust inference through soft knowledge aggregation of task sub-modules.

[0144] Specifically, the cross-subspace knowledge alignment method for continuous learning image classification proposed by the present invention has the following beneficial effects:

[0145] 1. The cross-subspace knowledge alignment and aggregation method provided by the present invention significantly improves the robustness of the model to incorrect task identifiers (task-ids). This feature is crucial for the class-incremental learning (CIL) method based on parameter-efficient fine-tuning (PEFT) because it can ensure the reliability and accuracy of the model even when the task identifiers are misleading;

[0146] 2. The present invention is particularly suitable for processing long-sequence tasks and can maintain the plasticity of the model even as the number of tasks increases. This means that no matter how many consecutive tasks are faced, the model can flexibly adapt to new tasks while effectively retaining the knowledge of the learned tasks, thus achieving better performance in the long-term learning process.

[0147] Through this design, the method of the present invention can not only cope with the challenge of incorrect task identifiers in the short term, but also adapt to the complex requirements of multi-task learning in the long term, providing an efficient, robust and flexible solution for continuous learning scenarios in practical applications.

[0148] The experimental process of the method of the embodiment of the present invention is described below.

[0149] The experiments were conducted on a single NVIDIA RTX 3090 GPU. ViT-B / 16 was used as the pre-trained backbone network, which was pre-trained on ImageNet-21K. The SGD optimizer with a momentum of 0.9 and an initial learning rate of 0.01 was used for training. Each task was trained for 10 epochs with a batch size of 110. The PEFT module (trainable prompt vectors and task adapters) was inserted into all transformer blocks to avoid searching. The prompt length was 4, and the intermediate dimension of the adapter was 64. For the experiments under the [10, 20, 50, 100]-split setting, the intermediate dimension of the adapter was set to [64, 32, 16, 8] to avoid additional inference overhead. In the knowledge alignment at the feature level, was set to 0.05, and in the knowledge alignment at the decision boundary level, and were 0.2 and 20 respectively. During the inference process, was set to 20 for the ImageNetA dataset and 10 for other settings, was set to 3.0.

[0150] The proposed method is evaluated on four common class-incremental learning (CIL) benchmark datasets, namely ImagenetR, ImagenetA, CIFAR100, and DomainNet. ImagenetR contains 30,000 images in 200 classes, which have the same class names as Imagenet-21K but belong to different domains. ImagenetA is a challenging dataset containing 7,500 images with difficult samples and significant class imbalance. CIFAR100 is a commonly used dataset in class-incremental learning, containing 60,000 images with a resolution of and a total of 100 classes. DomainNet is a cross-domain dataset containing images from 345 classes and 6 different domains. Referring to the settings of Cprompt, the present invention selects the 200 classes with the largest number of images for training and evaluation.

[0151] According to existing class-incremental learning methods based on parameter-efficient fine-tuning, the present invention adopts two common metrics to evaluate the model performance: the accuracy of all classes after learning the last task (Last-acc), and the average accuracy of all incremental tasks (Avg-acc). All experiments are conducted under three different random seeds, and the average performance and the maximum deviation are reported. The evaluation results are shown in Table 1.

[0152] Table 1 Comparison of the results of the method of the present invention and existing methods

[0153]

[0154]

[0155] For each method in Table 1, please refer to the related technology for understanding and no further explanation will be given here. The experiments show that the cross-subspace knowledge alignment and aggregation of the present invention is superior to existing class-incremental learning (CIL) methods based on parameter-efficient fine-tuning (PEFT) in terms of performance.

[0156] It should be noted that in the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0157] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0158] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included within the scope of protection of the present invention.

Claims

1. A cross-subspace knowledge alignment method for continuous learning image classification, characterized in that include: S1, obtaining a continuous learning training dataset for a category incremental learning scenario, wherein the continuous learning training dataset includes image datasets for multiple tasks of image classification, and the image datasets for all tasks have different categories and are disjoint; S2, building a network model for continuous learning image classification, the network model includes: a feature extractor formed by freezing the parameters of the pre-trained ViT network backbone network, L trainable task-shared prompt vectors, independent task submodules between trainable tasks, and a classifier constructed by a fully connected layer ; wherein the feature extractor comprises L transformer blocks, each hint vector is inserted into the original input of a transformer block respectively so that the L hint vectors and the feature extractor together constitute a task sharing network; each task submodule comprises L adapters corresponding to the L transformer blocks of the feature extractor; S3, determine the basic loss function of the network model, and construct the loss function of feature level alignment and the loss function of decision boundary level alignment, and use the sum of all loss functions to construct the total loss function, so as to construct a training framework for cross-subspace knowledge alignment; the basic loss function of the network model is expressed as: ; in, Represents the basic loss function of the network model; Indicates the current task Batches of image datasets size; Indicates the current task The first batch of the image dataset image data; Represents the softmax activation function commonly used in image classification; represents the cross entropy loss; the shared features extracted from the task sharing network are recorded as , will be added from the current task The features extracted by the task submodule are recorded as , Middle Shared Features , Middle Features ; represents a set of L prompt vectors; Indicates the task The task submodule; represents a feature extractor; express Through the classifier The characteristics of the latter; express Labels; express Through the classifier The characteristics of the latter; express Labels; S4, under the training framework, using the image data sets of the multiple tasks to train the network model in sequence until the value of the total loss function reaches convergence, thereby obtaining a trained network model; S5, inputting the test data set into the trained network model, realizing cross-subspace knowledge aggregation through multi-adapter hybrid reasoning guided by task confidence, thereby realizing the image classification task of continuous learning.

2. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 1, characterized in that: The continuous learning training data set is expressed as: ; Wherein, the continuous learning training data set include image datasets for tasks, is a natural number greater than 0; The image dataset for each task is represented as , The size is , express The image data, ; ; express The category label of Indicates the task The label space of different tasks does not intersect; , Indicates the height and width of the image data, Represents the number of channels of the image data.

3. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 2, characterized in that: Before inserting the L hint vectors, the feature extractor is used to: Input image data Divide patch blocks, and encode each patch block as dimensional vector to get the input image data Block coding features ;in, , and is a natural number greater than 0, The dimension representing the image features; Label the classification tokens and position embedding vector Added to the block encoding feature In the input image data Small block encoding representation of the image ,in, , , ; Encode the image in small blocks The feature extractor is processed by L transformer blocks, each of which consists of a multi-head self-attention MSA layer and a feed-forward network FFN with residual connections; The input of the converter block , the forward propagation process of the transformer block is expressed as: ; Indicates The output of the MSA layer in the transformer block is used as the The input of the FFN in the converter block, Indicates The input of the converter block; The classification token features output by L transformer blocks are used as input image data The image features output by the feature extractor are expressed as; , Represents a feature extractor.

4. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 3, characterized in that: After inserting L hint vectors, the task-sharing network is used to: For the Transformer blocks, the hint vector to be inserted is the same as the input of the transformer block Connect together as the input of the MSA layer of the transformer block, thereby changing the forward propagation process, and output shared features through the last layer of the task sharing network; Among them, the forward propagation process of the task sharing network is expressed by the formula: ; ; Among them, the set of L prompt vectors is expressed as , among which Hint vector Insert to In the original input of the transformer block, , represents the length of the inserted hint vector, Represents the dimension of image features, is a natural number greater than 0; Represents the task sharing network Output features of the MSA layer of the layer; Represents the task sharing network The input of the layer; The shared features output by the task sharing network are expressed as follows: .

5. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 4, characterized in that: Task The task submodule is represented as ;in, for The adapters, consisting of a downsampling matrix and an upsampling matrix composition, represents the intermediate feature dimension of the adapter, ; The forward propagation process is expressed as ; express Input; Each task submodule includes L adapters, corresponding to the L transformer blocks acting on the feature extractor, in the following manner: Task In the task submodule adapter , through the residual connection, it corresponds to the task The knowledge of In a FFN, the corresponding forward propagation process is expressed as: 。 6. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 5, characterized in that: The loss function of the feature level alignment is expressed as: ; in, represents the feature-level alignment loss function, which is used to measure the input batch The cross-subspace feature alignment loss, By shortening the distance between image features of the same type in different feature subspaces and increasing the distance between image features of different types in the same feature subspace, the features of samples of the same type in multiple subspaces are clustered together, thereby achieving semantic alignment of features of the same type. Represents positive samples in the feature space of different tasks; represents the positive sample set; represents the negative sample set; Indicates that a batch contains All features except It represents the temperature factor used for contrast learning; It means to find the union; represents the natural exponential function; Represents the cosine similarity function.

7. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 6, characterized in that: The loss function of the decision boundary level alignment is expressed as: ; in, represents the loss function for the decision boundary level alignment, used to train the task-adaptive classifier ; represents a unified feature set, including image features of all visible tasks in the current subspace; Indicates Adaptive classifier for each task; express Features in; express Adaptive classifier The characteristics of the latter; express .

8. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 7, characterized in that: After obtaining the trained network model, the method further includes: By summarizing the parameters of the adaptive classifiers of all tasks, a unified classifier suitable for the feature subspace of all tasks is constructed .

9. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 8, characterized in that: S5, inputting the test data set into the trained network model, realizing cross-subspace knowledge aggregation through multi-adapter hybrid reasoning guided by task confidence, thereby realizing the image classification reasoning task of continuous learning, including: For any test image data in the test data set, a classifier is obtained using the trained network model. The output and shared features of the classifier The output contains The confidence level of each category; is the total number of categories, which is a natural number greater than 0; The classifier The largest output elements are retained, and the remaining elements are set to infinity to obtain the updated classifier Output; Based on the updated classifier Output, obtain the confidence score that the shared feature belongs to each task; For each layer of the task sharing network, the confidence score of the shared feature belonging to each task is used as a routing weight, the adapter output features of the corresponding tasks of the layer are aggregated, and the adapter output features after aggregation of all tasks of the layer are integrated into the FFN of the layer as knowledge, and after passing through all layers, the aggregated features are output; The aggregated features are used to classify and the classifier , determine the predicted category of the test image data.

Citation Information

Patent Citations

  • Target domain data processing method, device and equipment based on context awareness

    CN117726887A

  • Visual question and answer method and system based on fine-grained adapter

    CN118607526A