Cross-subspace knowledge alignment method for continuous learning image classification
By proposing a cross-subspace knowledge alignment method in class incremental learning, using the alignment loss function at the feature level and decision boundary level, the feature subspace misalignment problem caused by task recognition error is solved, significantly improving the robustness and plasticity of the model, and achieving better image classification performance.
Patent Information
- Application Number
- CN202510421323.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In the category incremental learning scenario, due to the recognition error of the task recognizer, the test image may be projected into the wrong feature subspace, resulting in blurred classification results and lack of effective interaction and alignment.
A cross-subspace knowledge alignment method is proposed, and by constructing a network model, including a pre-trained ViT network backbone, a task-sharing prompt vector, an independent task submodule between tasks and a classifier with a fully connected layer. The loss function of feature-level alignment and decision-making boundary-level alignment is adopted to realize knowledge alignment and aggregation across subspaces.
It significantly improves the robustness of the model for wrong task identification, can ensure the reliability and accuracy of the model when task identification is misleading, is suitable for long-sequence tasks, maintains the plasticity of the model and achieves better performance in long-term learning.
Smart Images

Figure CN119942270A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of continuous learning image classification, and in particular relates to a cross-subspace knowledge alignment method for continuous learning image classification. Background Art
[0002] Continuous learning requires that the model can continuously learn from continuously incoming data streams to adapt to changing real-world scenarios. This is crucial for the deployment of artificial intelligence models in practical applications. However, continuous learning faces a major challenge: the model often encounters catastrophic forgetting in the process of learning new categories, that is, the knowledge of old categories is gradually lost. At present, there are two common scenarios for continuous learning: Task-Incremental Learning (TIL) and Class-Incremental Learning (CIL). In the TIL scenario, the task identifiers (task-ids) can be accessed during the reasoning phase, which is considered to be too simplified. The CIL scenario is more challenging because the task identifier is unknown during reasoning and the model needs to make inferences without clear task identifiers.
[0003] With the development of pre-trained models, methods based on parameter efficient fine-tuning (PEFT) have performed well in visual tasks. This type of method inserts lightweight modules for training on the basis of freezing the pre-trained backbone network, which can reduce the computational cost and training overhead while maintaining the generalization ability of the pre-trained model. In recent years, PEFT methods have been introduced into the field of continuous learning, especially in the category incremental learning scenario, showing strong performance. These methods capture task-specific knowledge and provide sufficient model plasticity for new tasks by assigning independent submodules to each task. Existing PEFT methods usually design task identifiers to identify task identifiers in category incremental learning, and select corresponding submodules to classify test images. However, due to the inevitable recognition errors of the task identifier, the test image may be projected into the wrong submodule subspace. In this case, the existing methods fail to establish effective interactions between the feature subspaces of different submodules due to the independent training of submodules and their classifiers, resulting in the problem of feature subspace misalignment. This problem manifests itself in two aspects: (1) There is inconsistency in the semantics of the same image features from the correct sub-module and the incorrect sub-module; (2) The decision boundary trained based on the correct task identity cannot effectively classify features in the incorrect subspace.
[0004] Based on the above analysis, the main issue to be considered in implementing this task is that due to the recognition error of the task identifier, the test image may be projected into the wrong feature subspace, resulting in ambiguous classification results. Therefore, how to effectively align and aggregate the feature subspaces of different submodules to improve the robustness of the model in the case of task recognition errors is the core issue to be solved. Summary of the invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a cross-subspace knowledge alignment method for continuous learning image classification. The technical problem to be solved by the present invention is achieved by the following technical solutions: A cross-subspace knowledge alignment method for continuous learning image classification, comprising: S1, obtaining a continuous learning training dataset for a category incremental learning scenario, wherein the continuous learning training dataset includes image datasets for multiple tasks of image classification, and the image datasets for all tasks have different categories and are disjoint; S2, building a network model for continuous learning image classification, the network model includes: a feature extractor formed by freezing the parameters of the pre-trained ViT network backbone network, L trainable task-shared prompt vectors, independent task submodules between trainable tasks, and a classifier constructed by a fully connected layer ; wherein the feature extractor comprises L transformer blocks, each hint vector is inserted into the original input of a transformer block respectively so that the L hint vectors and the feature extractor together constitute a task sharing network; each task submodule comprises L adapters corresponding to the L transformer blocks of the feature extractor; S3, determining the basic loss function of the network model, and constructing a feature-level aligned loss function and a decision boundary-level aligned loss function, and constructing a total loss function using the sum of all loss functions, thereby constructing a cross-subspace knowledge alignment training framework; S4, under the training framework, using the image data sets of the multiple tasks to train the network model in sequence until the value of the total loss function reaches convergence, thereby obtaining a trained network model; S5, inputting the test data set into the trained network model, realizing cross-subspace knowledge aggregation through multi-adapter hybrid reasoning guided by task confidence, thereby realizing the image classification task of continuous learning.
[0006] Beneficial effects of the present invention: 1. The present invention provides a cross-subspace knowledge alignment and aggregation method, which significantly improves the robustness of the model to incorrect task-ids. This feature is crucial for the category incremental learning (CIL) method based on parameter efficient fine-tuning (PEFT), because it can ensure the reliability and accuracy of the model even when the task-ids are misleading; 2. The present invention is particularly suitable for processing long sequence tasks, and can maintain the plasticity of the model even if the number of tasks continues to increase. This means that no matter how many consecutive tasks are faced, the model can flexibly adapt to new tasks while effectively retaining the knowledge of the learned tasks, thereby achieving better performance in the long-term learning process.
[0007] Through this design, the method of the present invention can not only cope with the challenge of task identification errors in the short term, but also adapt to the complex needs of multi-task learning in the long term, providing an efficient, robust and flexible solution for continuous learning scenarios in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A schematic diagram of a flow chart of a cross-subspace knowledge alignment method for continuous learning image classification provided by an embodiment of the present invention; Figure 2 The figure is a schematic diagram of the framework and principle of the network model of the embodiment of the present invention. DETAILED DESCRIPTION
[0009] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0010] In order to effectively align and aggregate the feature subspaces of different submodules and improve the robustness of the model in the case of task recognition errors, an embodiment of the present invention provides a cross-subspace knowledge alignment method for continuous learning image classification, such as Figure 1 As shown, the method may include the following steps: S1, obtaining a continuous learning training dataset for a category incremental learning scenario, wherein the continuous learning training dataset includes image datasets for multiple tasks of image classification, and the image datasets for all tasks have different categories and are disjoint; In the embodiment of the present invention, the continuous learning training data set is expressed as: ; Wherein, the continuous learning training data set include image datasets for tasks, is a natural number greater than 0; The image dataset for each task is represented as , The size is , express The image data, ; ; express The category label of Representation Task The label space of different tasks does not intersect; , Indicates the height and width of the image data, Indicates the number of channels of image data, , and are all natural numbers greater than 0.
[0011] In the embodiment of the present invention, All tasks are image classification tasks, but the number of classification categories can be different. For example, it can be a binary classification task or a multi-classification task.
[0012] S2, building a network model for continuous learning image classification, the network model includes: a feature extractor formed by freezing the parameters of the pre-trained ViT network backbone network, L trainable task-shared prompt vectors, independent task submodules between trainable tasks, and a classifier constructed by a fully connected layer ; wherein the feature extractor comprises L transformer blocks, each hint vector is inserted into the original input of a transformer block respectively so that the L hint vectors and the feature extractor together constitute a task sharing network; each task submodule comprises L adapters corresponding to the L transformer blocks of the feature extractor; See also Figure 2 The framework and principle diagram of the network model shown in the figure, the embodiment of the present invention is based on the ImageNet-21k dataset to pre-train the Vision Transformer (ViT) backbone network, and after the training is completed, the parameters are frozen to obtain the feature extractor; the feature extractor contains L transformer blocks, each of which consists of an MSA (Multi-head self-attention) layer and an FFN (Feed-Forward Neural Network) with residual connections. For simplicity, Figure 2 The complete structure of the L converter blocks is not illustrated in the figure, only MSA and FFN are given to illustrate the main composition of the converter blocks. Figure 2In the above, task-shared cues refer to the cue vectors shared between trainable tasks, which are implemented based on visual cue tuning; task-independent adapters refer to task-independent task submodules between trainable tasks. Each task corresponds to a task submodule; the classifier Refers to the classifier constructed by the fully connected layer .
[0013] In order to facilitate understanding of the structure and processing principle of the network model in the embodiment of the present invention, the feature extractor is first described.
[0014] Specifically, before inserting the L hint vectors, the feature extractor is used to: (1) Input image data Divide patch blocks, and encode each patch block as dimensional vector to get the input image data Block coding features ; Among them, the input image data , , Indicates the height and width of the image data, Indicates the number of channels of image data; and is a natural number greater than 0, The dimension representing the image features; (2) Mark the classification token and position embedding vector Added to the block encoding feature In the input image data Small block encoding representation of the image , in, , , (1); The above two steps can be understood as the pre-processing process of the feature extractor before the L transformer blocks, which can be implemented by a patch embedding layer (Patch layer), which is a learnable linear projection layer that transforms the input image data (i.e. input image) is evenly divided into Each block is named patch block. Each of the patch blocks will be encoded and mapped as dimensional vector, so that the input image data Transformed into block encoding features .
[0015] Classification Token Labeling is a learnable classification vector with initialization information when setting, Contains multiple elements, corresponding to the confidence of each category. Through the model learning process, it will eventually form the output of the feature extractor.
[0016] Position Embedding Vector express The relative position information of all elements in the .
[0017] The processing process of the Patch layer can be understood by referring to the relevant technology, and will not be explained in more detail here.
[0018] (3) Encode the image in small blocks Processed by L transformer blocks of the feature extractor, wherein each transformer block consists of a multi-head self-attention MSA layer and a feed-forward network FFN with residual connections; For the The input of the converter block , the forward propagation process of the transformer block is expressed as: (2); in, Indicates The output of the MSA layer in the transformer block is used as the The input of the FFN in the converter block, Indicates The input of the converter block; and Represent the processing of MSA layer and FFN respectively.
[0019] That is to say, for the A converter block with an input After being processed by its own MSA layer, the output , Enter its own FFN for processing and output As the The input of the converter block.
[0020] As can be seen from the previous text, the input to the first transformer block is the image block encoding representation ,Right now of , then the output of the first converter block, i.e. the input of the second converter block, is , and so on, the output of the Lth converter block is .
[0021] (4) The classification token features output after L transformer blocks are used as input image data The image features output by the feature extractor are expressed as: (3); in, represents a feature extractor; represents the output after L converter blocks are connected Extract the first element in the , that is, extract the .
[0022] The embodiment of the present invention develops a task sharing network based on Visual Prompt Tuning (VPT). Since visual prompts can inject knowledge as global operations before each transformer block, they are very suitable for capturing common knowledge of downstream tasks and have strong generalization ability.
[0023] Next, the L prompt vectors and the task sharing network formed by inserting the L prompt vectors into the feature extractor and the feature extractor are explained.
[0024] Specifically, the L cue vectors are a set of learnable vectors shared by a set of tasks, and their set is represented as , after inserting L hint vectors, the task-sharing network is used to: For the Transformer blocks, the hint vector to be inserted is the same as the input of the transformer block Connect together as the input of the MSA layer of the transformer block, thereby changing the forward propagation process, and output shared features through the last layer of the task sharing network; Among them, the forward propagation process of the task sharing network is expressed by the formula: ; (4); Among them, the set of L prompt vectors is expressed as , among which Hint vector Insert to In the original input of the transformer block, , represents the length of the inserted hint vector, Represents the dimension of image features, is a natural number greater than 0; Represents the task sharing network Output features of the MSA layer of the layer; Represents the task sharing network The input of the layer; The shared features output by the task sharing network are expressed as follows: (5); Compared with the separate feature extractor in formula (2), the input of the MSA layer in formula (4) is Based on this, an inserted hint vector is added , so that the output of the MSA layer becomes , the output It is still input into the corresponding FFN for processing. It can be understood that, from formula (2) to formula (4), Although both indicate The output of the converter block has actually changed.
[0025] Combining formulas (3) and (5), we can see that before inserting L hint vectors, the output of the feature extractor is , after inserting L hint vectors, the output of the feature extractor (which actually becomes a task-sharing network) is .
[0026] Since the task sharing network needs to adapt to all downstream tasks, it is necessary to reduce its forgetting phenomenon when training new tasks. To achieve this goal, the present invention updates along a direction orthogonal to the previous feature space. . It is updated in all consecutive tasks, encapsulating the common knowledge between them. With the help of this orthogonal optimization, a shared backbone network and powerful task identifier are built, which are specially tailored for downstream tasks.
[0027] Due to strict orthogonality constraints, the task sharing network is not sufficient to capture the fine-grained knowledge specific to each individual task, thus limiting the plasticity of the model. Therefore, the present invention assigns a task-specific task submodule to each task to capture its unique knowledge and generate more discriminative features. The task submodule of the present invention is designed based on the adapter architecture.
[0028] With a specific task For example, the task The task submodule is represented as ; in, for The adapter, acting on the feature extractor’s In the layers, due to the residual connection nature of adapter tuning, it can effectively enable the model to aggregate unique knowledge from different tasks independently.
[0029] A downsampling matrix and an upsampling matrix composition, represents the intermediate feature dimension of the adapter, ; The forward propagation process is expressed as: (6); express Input; express Activation function; In the embodiment of the present invention, each task submodule includes L adapters, corresponding to the L transformer blocks acting on the feature extractor, in the following manner: Task In the task submodule adapter , through the residual connection, it corresponds to the task The knowledge of In a FFN, the corresponding forward propagation process is expressed as: (7); See formula (4) , after adding the task submodule, it becomes formula (7), adding part.
[0030] S3, determining the basic loss function of the network model, and constructing a feature-level aligned loss function and a decision boundary-level aligned loss function, and constructing a total loss function using the sum of all loss functions, thereby constructing a cross-subspace knowledge alignment training framework; First, determine the basic loss function of the network model: In the During the task, a training dataset of size Batch (For simplicity, superscripts are omitted ), from the shared network and Extract features from task submodules and ,in , The output features of the network are then passed through a classifier constructed by a fully connected layer. ,for For classification problems, the dimension of the fully connected layer is . Network features and classifiers The commonly used cross entropy loss is used for training.
[0031] The basic loss function of the network model is expressed as: (8); in, Represents the basic loss function of the network model. Figure 2 In the expression ; Indicates the current task Batches of image datasets size; Represents the softmax activation function commonly used in image classification; represents the cross entropy loss; the shared features extracted from the task sharing network are recorded as , will be added from the current task The features extracted by the task submodule are recorded as , Middle Shared Features , Middle Features ; express Through the classifier The characteristics of the latter; express Labels; express Through the classifier The characteristics of the latter; express .
[0032] In view of the design of task-specific task submodules in the present invention, it is crucial to select the appropriate submodule by estimating the task ID through the task identifier. However, due to the limited performance of the task identifier, incorrect task IDs are inevitable. In order to enhance the robustness of the model to incorrect task IDs, a dual-level knowledge alignment (DKA) training method is proposed, which aligns task-specific knowledge at the feature and classifier levels.
[0033] (I) Feature level alignment: Generally speaking, during inference, it is only necessary to combine the feature extractor, the hint vector, and the task submodule corresponding to the test image to generate image features. However, since the task label is unknown during inference, it is very likely that the test image will be assigned to the task submodule corresponding to the wrong task label. Different task submodules will project the same image into different feature subspaces.
[0034] For a robust feature extractor that can generate meaningful features, even if the task ID is incorrect, the same image should achieve consistent semantics when projected into the correct and incorrect feature subspaces. To this end, the present invention unifies the semantics of features in different feature subspaces through feature level alignment. Specifically, when the same type of images pass through the network constructed by different task submodules, they should also output features that are as similar as possible.
[0035] The present invention proposes a cross-subspace feature alignment (CSFA) loss, which encourages the features of the same category in the current subspace to be aligned with the features in the previous subspace, that is, samples of the same category in the current task will output similar features after passing through the task submodule of the current task and the task submodule of the past task.
[0036] Because the training process of the network model is carried out in sequence by multiple tasks, that is, using Carry out corresponding task training; first for the current image dataset , extract these images in The features in the subspace ( ), that is, for the current task, its image dataset is passed through the current task submodule and the past task submodule to obtain image features, which are recorded as and During training, a set of features is randomly sampled in the prior subspace .
[0037] Given an input batch , for the label in the current subspace Features , whose positive set is constructed by features from a sampled set with the same ground truth, where At the same time, the features with different labels in the current subspace are regarded as negative samples, where .
[0038] That is, for images belonging to the same class, the present invention regards their features in the current subspace and the past subspace as positive sample pairs, and shortens their distance in the loss function. Conversely, for images belonging to different classes, their features in the current subspace are regarded as negative sample pairs, and shorten their distance in the loss function.
[0039] For input batch , its CSFA loss function, that is, the loss function of the feature level alignment, is expressed as: (9); in, represents the feature-level alignment loss function, which is used to measure the input batch The cross-subspace feature alignment loss, By shortening the distance between image features of the same type in different feature subspaces and increasing the distance between image features of different types in the same feature subspace, the features of samples of the same type in multiple subspaces are clustered together, thereby achieving semantic alignment of features of the same type. Represents positive samples in the feature space of different tasks; represents the positive sample set; represents the negative sample set; Indicates that a batch contains All features except It represents the temperature factor used for contrastive learning; It means to find the union; represents the natural exponential function; Represents the cosine similarity function.
[0040] CSFA loss achieves semantic alignment of similar features by shortening the distance between image features of the same type in different feature subspaces and increasing the distance between different images in the same feature subspace, and by clustering the features of samples of the same type in multiple subspaces.
[0041] (II) Decision boundary level alignment: Due to the inevitable subspace gaps, a classifier trained only with the correct task ID cannot handle features in the incorrect subspace. Decision boundary alignment aims to learn a sharp decision boundary for these incorrectly projected features. Specifically, in In the task, the present invention trains a task adaptive classifier , when the previous data , is projected into the current feature subspace, the task adaptive classifier tries to avoid ambiguous predictions. That is, each task will train its own task adaptive classifier, which is similar to the classifier constructed by the fully connected layer described above. Different from the previous one, the task adaptive classifier takes into account the image features of the current task and past tasks, so that the output prediction value of the task adaptive classifier will not be biased towards the current task.
[0042] To achieve this goal, the present invention first simulates the features of previous data in the current subspace, and then uses the simulated features to jointly train with the features of the current data. The present invention further decomposes the simulation process into three steps: 1. Model the affinity between current and previous data in a shared feature space through a task-sharing network; 2. Model the subspace shift in the current data between the current subspace and the shared subspace; 3. Shift the subspace of the current shared feature to the previous feature through the affinity matrix.
[0043] Since modeling affinity requires a prior data distribution, but no samples are saved for training, the present invention solves this problem through Gaussian sampling. After the task, for the of the tasks classes, and approximate their feature distribution in the shared feature space as a Gaussian distribution In the During each task, for each arriving batch, the shared features are distributed from the previous Mid-sample pseudo-representation , the size is , Represents the sampled pseudo representation set The Pseudo-features, express Combined with the current features and ,in Indicates shared features, express The category label of Indicates the first feature extracted by the task submodule of the current task indivual, express The category labels of the shared feature space are the affinity matrix between the previous data and the current data. It can be expressed as follows: (10); in, Represents the affinity matrix Middle Line Elements of a column; represents the temperature coefficient; Pseudo-representation set Zhongli Recent Features.
[0044] The subspace offset of the current data between the current subspace and the shared feature space can be simply obtained by using the specific feature Shared features with current data Then, the subspace is shifted from the current feature Transfer to sampling pseudo features . Simulation characteristics It is expressed as follows: (11); in, Represents a simulated feature.
[0045] Then we get a unified feature set , which includes the image features of all visible tasks in the current subspace. Use it to train a task-adaptive classifier , the loss function of the decision boundary level alignment is expressed as: (12); in, represents the loss function for the decision boundary level alignment, used to train the task-adaptive classifier ; represents a unified feature set, including image features of all visible tasks in the current subspace; Indicates Adaptive classifier for each task; express Features in; express Adaptive classifier The characteristics of the latter; express .
[0046] Therefore, for each task, the corresponding adaptive classifier is trained respectively. After all the training is completed, in order to reduce the computational overhead of inference, after the trained network model is obtained, the method further includes: By summarizing the parameters of the adaptive classifiers of all tasks, a unified classifier suitable for the feature subspace of all tasks is constructed .
[0047] The present invention summarizes the parameters of all task adaptive classifiers, namely , thus building a unified classifier ; The parameters of the adaptive classifiers for all tasks are aggregated using the parameter aggregation formula. Task No. The parameters corresponding to the categories are , the parameter aggregation formula is as follows: (13); in, It is a unified classifier Middle The weight of each category; It is a unified classifier Middle Bias of categories; is a task-adaptive classifier Middle The weight of each category; is a task-adaptive classifier Middle Bias of categories; A variable representing the number of the task, represents the number of the current task. This aggregation strategy enables the unified classifier to distinguish features projected into various subspaces, even in the face of misleading task IDs.
[0048] The basic loss function, the feature-level alignment loss function, and the decision boundary-level alignment loss function are summed to obtain the total loss function, which is expressed as: (14); The entire cross-subspace knowledge alignment training framework is trained by minimizing the total loss function.
[0049] For simplicity, the weighting parameters of all loss terms are set to 1.0. In each training session, the training framework of the present invention can be optimized in an end-to-end manner. This approach can keep the decision boundaries of the feature space corresponding to different task submodules consistent, so that when the task label of the test image is unknown and the wrong task submodule is selected, reliable classification results can still be obtained.
[0050] In summary, the present invention constructs a "cross-subspace knowledge alignment" training framework, which performs double-layer alignment of features and decision boundaries respectively. At the feature level, features of the same category are aligned in different feature subspaces, so that the distribution of features of the same category in cross-task submodules tends to be consistent; at the classifier level, a task adaptive classifier is trained separately for each task, which can classify features of all learned categories, even if these features are incorrectly projected into the feature subspace of the current task.
[0051] S4, under the training framework, using the image data sets of the multiple tasks to train the network model in sequence until the value of the total loss function reaches convergence, thereby obtaining a trained network model; The network model is trained according to the training framework in S2 and S3 and the total loss function therein, and its parameters are continuously updated to achieve the generation of the optimal network model, that is, to obtain the trained network model.
[0052] For details, please refer to the above training framework and combine it with the training process of the general continuous learning image classification model. I will not explain it in detail here.
[0053] S5, inputting the test data set into the trained network model, realizing cross-subspace knowledge aggregation through multi-adapter hybrid reasoning guided by task confidence, thereby realizing the image classification task of continuous learning.
[0054] S5 may include the following steps: S51, for any test image data in the test data set, obtain a classifier using the trained network model The output and shared features of the classifier The output contains The confidence level of each category; is the total number of categories, which is a natural number greater than 0; The test dataset may include test image datasets for multiple tasks.
[0055] In the experiment of the embodiment of the present invention, for ease of understanding, the test data set is represented as ,include The test image dataset for the task; The test image dataset for each task is represented as , The size is , express The Test image data, ; , express The category label of Representation Task The label space of different tasks does not intersect. , Indicates the height and width of the image data, Represents the number of channels of the image data.
[0056] The test data set is input into the trained network model to perform test and image category prediction. Given the task specificity of the architecture of the present invention, it is critical to aggregate task-specific knowledge during reasoning without the need for a task ID, and to project the test image (i.e., test image data) into an aligned feature space. To achieve this goal, the present invention utilizes task-shared knowledge from the task identifier to aggregate task-specific knowledge associated with each test image.
[0057] Given a test image data and their shared characteristics , before the softmax classifier The output is expressed as , and No. The elements are represented as . is the total number of visible categories.
[0058] S52, the classifier The largest output elements are retained, and the remaining elements are set to infinity to obtain the updated classifier Output; In order to extract relevant information from each task, the present invention only retains the classes that are semantically related to the test image and discards other classes. The largest elements corresponding to the categories most semantically similar to the test image: Specifically, you can In Sort the elements in descending order and select the one that comes first. elements, and then This elements are retained, and the remaining elements are set to infinity, and the updated classifier is obtained. Output.
[0059] The above process can be expressed as: (15); in, express The update result of express Top- The largest element.
[0060] S53, based on the updated classifier Output, obtain the confidence score that the shared feature belongs to each task; Specifically, it is concluded Belong to the task The confidence score of ,in Represents the amount of semantic information related to the input contained in each task: (16); in, express Belong to the task probability; is the temperature that controls the smoothness of the task confidence scores.
[0061] S54, for each layer of the task sharing network, using the confidence score of the shared feature belonging to each task as a routing weight, aggregating the adapter output features of the tasks corresponding to the layer, integrating the adapter output features after aggregation of all tasks in the layer as knowledge into the FFN of the layer, and outputting the aggregated features after passing through all layers; Confidence score is used as routing weights to aggregate task-specific intermediate features to form a hybrid adapter (MoA) architecture. Given the input of the network model is , No. The input of each converter block is , the forward propagation process is as follows: (17); Compared with formula (7), formula (17) converts the individual adapter processing result of the transformer block into the weighted result of the adapter result of the transformer block and the corresponding confidence score on all tasks.
[0062] Compared with assigning a single task to each test image, this mixture of adapters (MoA) structure avoids overconfident aggregation and further enhances the model's robustness to misleading task IDs.
[0063] The aggregated features after the Lth block are used for the final classification, expressed as .
[0064] S55, the aggregated features are classified using the unified classifier and the classifier , determine the predicted category of the test image data.
[0065] Final prediction By unified classifier and classifier Jointly determine: (18); in, Indicates the predicted category; The present invention aims to enhance the robustness of the model against misleading task IDs, including a knowledge alignment training strategy and a knowledge aggregation reasoning scheme. Through the double-layer alignment of feature semantics and decision boundaries and task-adaptive classifier aggregation, the subspace misalignment problem caused by task identification errors is alleviated, and more robust category incremental learning (CIL) is achieved. Specifically, cross-subspace knowledge alignment and aggregation ensure the plasticity of the model through dedicated task-specific submodules while learning task identifiers, and a dual knowledge alignment training framework is adopted. Dual knowledge alignment aligns feature semantics and decision boundaries of different subspaces at the feature and classifier decision boundary levels, thereby overcoming the fuzzy decision problem caused by subspace misalignment. In addition, in order to avoid overconfidence in misleading task IDs during the reasoning process, a task confidence-guided adapter hybrid powerful reasoning mechanism is proposed, which achieves more robust reasoning through soft knowledge aggregation of task submodules.
[0066] Specifically, the cross-subspace knowledge alignment method for continuous learning image classification proposed in the present invention has the following beneficial effects: 1. The cross-subspace knowledge alignment and aggregation method provided by the present invention significantly improves the robustness of the model to incorrect task-ids. This feature is crucial for the category incremental learning (CIL) method based on parameter efficient fine-tuning (PEFT), because it can ensure the reliability and accuracy of the model even when the task-ids are misleading; 2. The present invention is particularly suitable for processing long sequence tasks, and can maintain the plasticity of the model even if the number of tasks continues to increase. This means that no matter how many consecutive tasks are faced, the model can flexibly adapt to new tasks while effectively retaining the knowledge of the learned tasks, thereby achieving better performance in the long-term learning process.
[0067] Through this design, the method of the present invention can not only cope with the challenge of task identification errors in the short term, but also adapt to the complex needs of multi-task learning in the long term, providing an efficient, robust and flexible solution for continuous learning scenarios in practical applications.
[0068] The following is a description of the experimental process of the method according to the embodiment of the present invention.
[0069] The experiments were conducted on a single NVIDIA RTX 3090 GPU. ViT-B / 16 was used as the pre-trained backbone network, which was pre-trained on ImageNet-21K. Training was performed using an SGD optimizer with a momentum of 0.9 and an initial learning rate of 0.01. Each task was trained for 10 epochs with a batch size of 110. The PEFT module (trainable hint vector and task adapter) was inserted into all transformer blocks to avoid search. The hint length is 4 and the intermediate dimension of the adapter is 64. For experiments under the [10, 20, 50, 100]-split setting, the intermediate dimension of the adapter is set to [64, 32, 16, 8] to avoid additional inference overhead. In feature-level knowledge alignment, Set to 0.05, the knowledge alignment at the decision boundary level and are 0.2 and 20 respectively. During the inference process, For the ImageNetA dataset, it is set to 20, and for the others, it is set to 10. Set to 3.0.
[0070] The proposed method is evaluated on four common category incremental learning (CIL) benchmark datasets, namely ImagenetR, ImagenetA, CIFAR100 and DomainNet. ImagenetR contains 30,000 images in 200 categories, which have the same category names as Imagenet-21K but belong to different domains. ImagenetA is a challenging dataset containing 7,500 images with difficult samples and significant category imbalance. CIFAR100 is a commonly used dataset in category incremental learning, containing 60,000 images with a resolution of There are 100 categories of images in total. DomainNet is a cross-domain dataset containing images from 345 categories and 6 different domains. Referring to the settings of Cprompt, the present invention selects the 200 categories with the largest number of images for training and evaluation.
[0071] Based on the existing category incremental learning method based on efficient parameter fine-tuning, this paper uses two common indicators to evaluate model performance: the accuracy of all categories after learning the last task (Last-acc), and the average accuracy of all incremental tasks (Avg-acc). All experiments were conducted under three different random seeds, and the average performance and maximum deviation were reported. The evaluation results are shown in Table 1.
[0072] Table 1 Comparison of the results of the method of the present invention and the existing method
[0073]
[0074] Please refer to the related art for understanding the methods in Table 1, and no further explanation is given here. Experiments show that the cross-subspace knowledge alignment and aggregation of the present invention outperforms the existing category incremental learning (CIL) method based on efficient parameter fine tuning (PEFT) in performance.
[0075] It should be noted that, in the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.
[0076] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0077] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.
[0078] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A cross-subspace knowledge alignment method for continuous learning image classification, characterized in that include: S1, obtaining a continuous learning training dataset for a category incremental learning scenario, wherein the continuous learning training dataset includes image datasets for multiple tasks of image classification, and the image datasets for all tasks have different categories and are disjoint; S2, building a network model for continuous learning image classification, the network model includes: a feature extractor formed by freezing the parameters of the pre-trained ViT network backbone network, L trainable task-shared prompt vectors, independent task submodules between trainable tasks, and a classifier constructed by a fully connected layer ; wherein the feature extractor comprises L transformer blocks, each hint vector is inserted into the original input of a transformer block respectively so that the L hint vectors and the feature extractor together constitute a task sharing network; each task submodule comprises L adapters corresponding to the L transformer blocks of the feature extractor; S3, determining the basic loss function of the network model, and constructing a feature-level aligned loss function and a decision boundary-level aligned loss function, and constructing a total loss function using the sum of all loss functions, thereby constructing a cross-subspace knowledge alignment training framework; S4, under the training framework, using the image data sets of the multiple tasks to train the network model in sequence until the value of the total loss function reaches convergence, thereby obtaining a trained network model; S5, inputting the test data set into the trained network model, realizing cross-subspace knowledge aggregation through multi-adapter hybrid reasoning guided by task confidence, thereby realizing the image classification task of continuous learning.
2. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 1, characterized in that: The continuous learning training data set is expressed as: ; Wherein, the continuous learning training data set include image datasets for tasks, is a natural number greater than 0; The image dataset for each task is represented as , The size is , express The image data, ; ; express The category label of Indicates the task The label space of different tasks does not intersect; , Indicates the height and width of the image data, Represents the number of channels of the image data.
3. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 2, characterized in that: Before inserting the L hint vectors, the feature extractor is used to: Input image data Divide patch blocks, and encode each patch block as dimensional vector to get the input image data Block coding features ;in, , and is a natural number greater than 0, The dimension representing the image features; Label the classification tokens and position embedding vector Added to the block encoding feature In the input image data Small block encoding representation of the image ,in, , , ; Encode the image in small blocks The feature extractor is processed by L transformer blocks, each of which consists of a multi-head self-attention MSA layer and a feed-forward network FFN with residual connections; The input of the converter block , the forward propagation process of the transformer block is expressed as: ; Indicates The output of the MSA layer in the transformer block is used as the The input of the FFN in the converter block, Indicates The input of the converter block; The classification token features output by L transformer blocks are used as input image data The image features output by the feature extractor are expressed as; , Represents a feature extractor.
4. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 3, characterized in that: After inserting L hint vectors, the task-sharing network is used to: For the Transformer blocks, the hint vector to be inserted is the same as the input of the transformer block Connect together as the input of the MSA layer of the transformer block, thereby changing the forward propagation process, and output shared features through the last layer of the task sharing network; Among them, the forward propagation process of the task sharing network is expressed by the formula: ; ; Among them, the set of L prompt vectors is expressed as , among which Hint vector Insert to In the original input of the transformer block, , represents the length of the inserted hint vector, Represents the dimension of image features, is a natural number greater than 0; Represents the task sharing network Output features of the MSA layer of the layer; Represents the task sharing network The input of the layer; The shared features output by the task sharing network are expressed as follows: .
5. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 4, characterized in that: Task The task submodule is represented as ;in, for The adapters, consisting of a downsampling matrix and an upsampling matrix composition, represents the intermediate feature dimension of the adapter, ; The forward propagation process is expressed as ; express Input; Each task submodule includes L adapters, corresponding to the L transformer blocks acting on the feature extractor, in the following manner: Task In the task submodule adapter , through the residual connection, it corresponds to the task The knowledge of In a FFN, the corresponding forward propagation process is expressed as: 。 6. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 5, characterized in that: The basic loss function of the network model is expressed as: ; in, Represents the basic loss function of the network model; Indicates the current task Batches of image datasets size; Represents the softmax activation function commonly used in image classification; represents the cross entropy loss; the shared features extracted from the task sharing network are recorded as , will be added from the current task The features extracted by the task submodule are recorded as , Middle Shared Features , Middle Features ; express Through the classifier The characteristics of the latter; express Labels; express Through the classifier The characteristics of the latter; express .
7. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 6, characterized in that: The loss function of the feature level alignment is expressed as: ; in, represents the feature-level alignment loss function, which is used to measure the input batch The cross-subspace feature alignment loss, By shortening the distance between image features of the same type in different feature subspaces and increasing the distance between image features of different types in the same feature subspace, the features of samples of the same type in multiple subspaces are clustered together, thereby achieving semantic alignment of features of the same type. Represents positive samples in the feature space of different tasks; represents the positive sample set; represents the negative sample set; Indicates that a batch contains All features except It represents the temperature factor used for contrastive learning; It means to find the union; represents the natural exponential function; Represents the cosine similarity function.
8. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 7, characterized in that: The loss function of the decision boundary level alignment is expressed as: ; in, represents the loss function for the decision boundary level alignment, used to train the task-adaptive classifier ; represents a unified feature set, including image features of all visible tasks in the current subspace; Indicates Adaptive classifier for each task; express Features in; express Adaptive classifier The characteristics of the latter; express .
9. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 8, characterized in that: After obtaining the trained network model, the method further includes: By summarizing the parameters of the adaptive classifiers of all tasks, a unified classifier suitable for the feature subspace of all tasks is constructed .
10. The cross-subspace knowledge alignment method for continuous learning image classification according to claim 9, characterized in that: S5, inputting the test data set into the trained network model, realizing cross-subspace knowledge aggregation through multi-adapter hybrid reasoning guided by task confidence, thereby realizing the image classification reasoning task of continuous learning, including: For any test image data in the test data set, a classifier is obtained using the trained network model. The output and shared features of the classifier The output contains The confidence level of each category; is the total number of categories, which is a natural number greater than 0; The classifier The largest output elements are retained, and the remaining elements are set to infinity to obtain the updated classifier Output; Based on the updated classifier Output, obtain the confidence score that the shared feature belongs to each task; For each layer of the task sharing network, the confidence score of the shared feature belonging to each task is used as a routing weight, the adapter output features of the corresponding tasks of the layer are aggregated, and the adapter output features after aggregation of all tasks of the layer are integrated into the FFN of the layer as knowledge, and after passing through all layers, the aggregated features are output; The aggregated features are used to classify and the classifier , determine the predicted category of the test image data.
Citation Information
Patent Citations
Target domain data processing method, device and equipment based on context awareness
CN117726887A
Visual question and answer method and system based on fine-grained adapter
CN118607526A
Air-ground cross-platform target re-identification method based on semantic alignment and prompt learning
CN119007241A
Unified framework for vision prompt tuning
US20240378870A1
Brain-computer information fusion classification method and system for shared subspace learning
WO2023173804A1
Cited By
HRRP signal feature recognition method, device and equipment based on time-frequency domain fusion
CN121410671A
Hrrp signal feature recognition method, device and equipment based on time-frequency domain fusion
CN121410671B