Open vocabulary multi-task image classification method based on continuous learning
By fine-tuning the image encoder and introducing a guided attention module, combined with random projection and class prototype modeling, the catastrophic forgetting problem in incremental learning of multi-domain tasks is solved, the accuracy and stability of image classification are improved, and continuous learning capabilities in multi-task scenarios are achieved.
Patent Information
- Application Number
- CN202511096243.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing continuous learning image classification methods have difficulty in fully utilizing existing knowledge when faced with incremental learning scenarios for multi-domain tasks, resulting in catastrophic forgetting, and their classification accuracy and stability are insufficient in multi-task scenarios.
By fine-tuning the image encoder to enhance feature extraction capabilities, introducing a guided attention module to achieve deep fusion of image and text features, and adopting a random projection and prototype modeling mechanism, combined with a multi-loss joint training strategy, it alleviates the catastrophic forgetting problem and improves the model's cross-task generalization ability.
Maintain high classification accuracy and stability in multi-task scenarios, enhance the model's semantic representation and category discrimination capabilities, reduce forgetting problems, and improve the model's multi-task image classification performance in open environments.
Smart Images

Figure CN120599385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an open vocabulary multi-task image classification method based on continuous learning. Background Art
[0002] While machine learning models have demonstrated exceptional performance in certain single tasks in recent years, these models often rely on training in static environments and lack the ability to dynamically adjust their behavior over time. Consequently, retraining is often required whenever new data becomes available, an approach that struggles to adapt to the ever-changing data streams. To address this challenge, there is an urgent need to build systems with continuous adaptability that can continuously learn and accumulate knowledge in a constantly changing environment.
[0003] With the development of technology, research on image classification based on continuous learning has made significant progress. However, existing continuous learning image classification methods still have certain limitations when facing incremental learning scenarios for multi-domain tasks. Multi-domain incremental learning requires the model to be able to continuously learn across multiple image domains. Different tasks often have significant distribution differences, including changes in image style, background complexity, and category semantics. Such inter-domain differences make it difficult for the model to fully utilize existing knowledge when learning new tasks, thereby exacerbating catastrophic forgetting. Summary of the Invention
[0004] The purpose of the present invention is to enhance the feature extraction capability of the classification model by fine-tuning the image encoder, introduce a guided attention module to achieve deep fusion of image and text features, and improve the recognition ability of key semantic features. To this end, an open vocabulary multi-task image classification method based on continuous learning is provided.
[0005] In order to achieve the above-mentioned object of the invention, the embodiment of the present invention provides the following technical solutions:
[0006] The open vocabulary multi-task image classification method based on continuous learning includes the following steps:
[0007] Step 1: Obtain original image data, pre-process the original image data to obtain corresponding text information, input the text information into a text encoder, and obtain text features;
[0008] Step 2: Input the original image data into the image encoder to obtain image features;
[0009] Step 3: Input the text features and image features into the guided attention module, perform weighted integration on the image features, and obtain a multimodal feature representation that is semantically consistent with the text features;
[0010] Step 4: Input the multimodal features into the random projection module for random projection, and obtain the activated features through the nonlinear activation function; take the mean of the activated features of each category through the prediction module to generate the class prototype vector, and input the nonlinear activation function into the Gram matrix to obtain the category of the original image data.
[0011] Compared with the existing technology, the beneficial effects of the present invention are as follows: the present invention combines the continuous learning mechanism with the open vocabulary image recognition method, and can maintain high classification accuracy and stability in multi-task scenarios where there are significant domain differences between tasks. The feature extraction ability of the classification model is enhanced by fine-tuning the image encoder, and the guided attention module is introduced to achieve deep fusion of image and text features, thereby improving the recognition ability of key semantic features. In addition, the random projection and prototype modeling mechanism is adopted to perform structured modeling and aggregation of multimodal features, so that the model has stronger semantic representation ability and category discrimination ability when facing open vocabulary classification tasks. At the same time, the multi-loss joint training strategy is adopted to effectively alleviate the catastrophic forgetting problem, so that the model has stronger cross-task generalization ability and continuous adaptability, thereby improving the overall performance of multi-task image classification in an open environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 This is a schematic diagram of the classification model processing process according to an embodiment of the present invention;
[0014] Figure 2 Schematic diagram of the processing process of the random projection module and the prediction module according to an embodiment of the present invention. DETAILED DESCRIPTION
[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0016] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and are not to be understood as indicating or implying relative importance, or implying any actual relationship or order between these entities or operations. In addition, the terms "connected" and "connected" can refer to direct connection between elements or indirect connection via other elements.
[0017] Example 1:
[0018] The present invention is implemented by the following technical solution: an open vocabulary multi-task image classification method based on continuous learning, Figure 1 The classification model shown in the figure is implemented, and the classification model includes a text encoder, an image encoder, a guided attention module, a random projection module, and a prediction module. This method includes the following steps:
[0019] Step 1: obtain the original image data, preprocess the original image data to obtain the corresponding text information, input the text information into the text encoder, and obtain text features.
[0020] A large amount of raw image data suitable for multi-task classification is obtained from a variety of public datasets. Each public dataset represents a classification task and has different image domains and category labels. After obtaining the raw image data, it is subjected to standardized preprocessing, including resizing, normalization, and image enhancement, to improve the training stability and generalization ability of the classification model. For each image, corresponding text information is constructed for subsequent multimodal feature alignment. This text information is automatically generated based on predefined templates, and different predefined templates are corresponding to different public datasets. For example, for the Food101 dataset, the predefined template constructed is "a photo of {}, a type of food", where {} represents the specific food name corresponding to the image; for the StanfordCars dataset, the predefined template constructed is "a photo of a {}, a type of car", where {} represents the specific car model name corresponding to the image.
[0021] A text encoder is used to obtain text features. For example, the text information of "a photo of {}, a type of food" is first segmented and converted into a token sequence of length T. This is then mapped into an embedding vector sequence, and then a learnable positional encoding is added to form the input sequence:
[0022] Among them, W 0represents the input sequence; w t represents the embedding vector sequence of the tth word; p t represents the positional encoding corresponding to the t-th embedding vector sequence; t=1,2,...,T.
[0023] Input sequence W 0 The text is then fed into a multi-layer Transformer encoder (i.e., a text encoder) for context modeling. Each layer includes a multi-head self-attention mechanism and a feedforward network. Residual connections and layer normalization are used to ensure the stability of classification model training. Finally, the multi-layer Transformer encoder outputs a special token [EOS] at the end of the sentence, representing the text features that aggregate the text information of the entire sentence:
[0024] Among them, f text Represents text features; TransformerEncoder represents the operation of the text encoder; ( ) [EOS] Indicates the special token at the end of the sentence. The obtained text feature f text It is located in the same space as the image features and can be used for similarity calculation, attention interaction and joint classification with the image features. This embedding method not only improves the semantic consistency between images and text, but also enhances the classification model's ability to capture semantic details.
[0025] Step 2: Input the original image data into the image encoder to obtain image features;
[0026] The original image data As the input of the image encoder, H, W, and C represent the height, width, and number of channels of the original image data, respectively. The image encoder includes L stacked Transformer encoding blocks. In this embodiment, the first Transformer encoding block is used as the frozen image encoder, and the second Transformer encoding block is used as the fine-tuned image encoder.
[0027] Before the original image data is input into the first Transformer encoding block, in order to adapt the Transformer architecture to the processing request of sequence data, it is necessary to first divide the input original image data I into non-overlapping image blocks (patches) of size P×P, so as to obtain a total of image blocks, flatten each image block into a vector to form the initial patch sequence {x1,x2,...,x N}, where each , i=1,2,...,N; is a set of real numbers. Then, each image patch is projected through a learnable linear mapping layer (Patch Embedding) and mapped to a unified feature dimension space. , get the embedding vector corresponding to the image block:
[0028] in, represents the embedding vector corresponding to the i-th image block; i=1,2,...,N; x i The vector representing the i-th image block; represents the linear projection matrix; Represents the two-dimensional position encoding vector corresponding to the i-th image block, which is used to retain the spatial structure information of each image block. The introduction of position encoding avoids the loss of position information of the patch sequence in sequence processing. D is the feature dimension.
[0029] The embedding vector sequence of N image patches First, it is fed into the first Transformer encoder block. Each Transformer encoder block contains a multi-head self-attention module and a feedforward network, with residual connections and layer normalization operations:
[0030] Among them, Z l represents the output of the lth Transformer encoding block; Z l-1 represents the output of the l-1th Transformer encoding block; FFN represents the operation of the feedforward network; MHSA represents the operation of the multi-head self-attention module. For example:
[0031] Until the output Z of the Lth Transformer encoding block is obtained L .
[0032] In the L stacked Transformer encoding blocks, the original image data I first passes through the frozen image encoder (composed of the frozen Transformer encoding blocks in the previous part) to extract the intermediate features:
[0033] in, is the intermediate feature; x I Represents the original image data; Represents a frozen image encoder, whose parameters do not participate in training and updating.
[0034] Next, the intermediate feature z intInput the fine-tuned image encoder (composed of the Transformer encoding block fine-tuned in the latter part) to obtain the final image features:
[0035] Among them, v represents the image feature; represents a fine-tuned image encoder; Represents the fine-tuning parameters.
[0036] During the training process, the total loss function L is defined total , this scheme aims to fine-tune the parameters of the image encoder Perform gradient updates:
[0037] Among them, for the fine-tuned image encoder, represents the fine-tuning parameters before updating; represents the updated fine-tuning parameters; represents the learning rate; is the current total loss function L total About fine-tuning parameters The gradient of The direction where the loss changes fastest. Due to the frozen image encoder It does not participate in gradient calculation, so it ensures that the original information will not be overwritten by the new task.
[0038] The frozen image encoder leverages the powerful feature extraction capabilities of the pre-trained model to provide stable and universal intermediate image features without updating its parameters. This freezing operation effectively prevents the destruction of existing knowledge during incremental learning, thereby reducing the risk of catastrophic forgetting, ensuring that the model's performance on previous tasks does not significantly degrade, and enhancing overall stability. The fine-tuned image encoder, based on the frozen image encoder, only updates the parameters of specific high-level modules. This design enables the model to adapt to the image style and semantic differences between different tasks while maintaining the universality of the underlying features, improving the model's expressive and discriminative capabilities for the current task, and enhancing the model's plasticity and generalization capabilities.
[0039] For image encoders, through the coordinated cooperation of freezing and fine-tuning, rapid adaptation to new tasks is achieved while maintaining the stability of the original knowledge, effectively improving the accuracy, robustness and task transferability of image classification in continuous learning scenarios.
[0040] like Figure 1As shown in the figure, when the original image data passes through the image encoder, it must first be processed by the frozen image encoder and then by the fine-tuned image encoder. Therefore, distillation constraints are introduced to maintain the representational consistency of the output image features with the reference image features. Specifically, the reference images come from the large-scale general image recognition dataset ImageNet. This dataset has the characteristics of diverse categories and wide distribution of visual semantics, which can provide stable and high-quality visual reference features for the model. The image encoder used for the reference images is the initial image encoder. Unlike the image encoder whose parameters are fine-tuned in some structures during training, the initial image encoder remains completely frozen and does not participate in fine-tuning. Its parameters are consistent with the pre-trained model, ensuring that the extracted reference features have good versatility and stability.
[0041] In this method, to enhance the image encoder's ability to perceive the semantics of the current task, portions of the network structure are fine-tuned. However, while this process improves the model's adaptability to specific tasks, it can cause significant drift in the image feature embedding space, causing their distribution to deviate from the visual representation space of the original pre-trained model, thereby impacting the model's generalization performance. To address this issue, a distillation constraint function based on reference image features is introduced. This effectively prevents the model from deviating excessively from the original feature distribution during fine-tuning, avoids the loss of general visual knowledge, and improves stability during continuous learning and task transfer. Furthermore, the attention module is guided to rely on the spatial structure of image features to generate keys and values during image-text fusion. Drastic changes in image features across different tasks can severely impact the attention mechanism's ability to focus on key semantic regions. Therefore, the distillation constraint function not only provides stability constraints at the feature representation level but also plays a key regulatory role in the multimodal fusion process, significantly enhancing the model's robustness and generalization in complex scenarios such as continuous learning and multi-task recognition.
[0042] In step 3, the text features and image features are input into the guided attention module to perform weighted integration on the image features to obtain a multimodal feature representation that is semantically consistent with the text features.
[0043] After inputting text and image features into the guided attention module, the text features are mapped into query vectors, and the image features are mapped into key and value vectors. The attention mechanism then fuses the text and image information. Specifically, the guided attention module dynamically guides the attention weighting of image features using text features, thereby improving the classification model's ability to perceive key semantic regions in multi-task images.
[0044] Assume that the image feature output by the image encoder is represented as , the text features output by the text encoder are , where B represents the batch size, D is the feature dimension, H and W are the length and width of the original image data respectively. First, the text feature f text Mapped to query vector Q, the image feature v is mapped to key vector K and value vector V:
[0045] in, is a learnable parameter matrix; Flatten() represents the architecture image feature flattened into a two-dimensional tensor of shape B×(H×W)×D.
[0046] Then, expand the text feature dimension so that , dot product with the key vector K of image features to obtain the attention score S:
[0047] Where d represents the dimension of the query vector Q and the key vector K.
[0048] Next, the attention score S is normalized by Softmax to obtain the attention weight A:
[0049] Use the attention weight A to perform weighted summation on the value vector V to obtain the fused multimodal feature F attn :
[0050] Finally, after removing redundant dimensions, the final multimodal features are obtained:
[0051] Among them, f attn Indicates multimodal features with consistent semantics between image features and text features; Squeeze( ) represents the operation of removing redundant dimensions.
[0052] Step 4: Input the multimodal features into the random projection module for random projection, and obtain the activated features through the nonlinear activation function; take the mean of the activated features of each category through the prediction module to generate the class prototype vector, and input the nonlinear activation function into the Gram matrix to obtain the category of the original image data.
[0053] like Figure 2 As shown, the random projection module performs attnPerforming random projections and mapping them to multiple low-dimensional projection subspaces can effectively improve the generalization ability and continuous learning ability of the classification model between tasks. Specifically, let the multimodal features output by the guided attention module be f attn , first use R linearly invariant random projection matrices Map it:
[0054] Among them, z r represents the activation feature under the r-th projection subspace; P r represents the rth projection matrix, R is the number of projection matrices; represents a nonlinear activation function (such as ReLU or GELU); M represents the projection dimension, , for example, D=512, M=10000.
[0055] For category c, the prediction module collects z r Activation feature representation in training samples , and take the average value to form the class prototype vector:
[0056] Among them, p c,r Represents the class prototype vector of the c-th category; represents the i-th activated feature in category c under the r-th projection subspace, i=1,2,...,N c , N c Represents the total number of features in category c. For example, if there are 1000 training samples in category c, then the category is considered to have 1000 features, and i is the feature index.
[0057] Next, we activate the feature z by calculating r With each class prototype vector p c,r The inner product similarity between them is used to construct the Gram matrix to obtain the score of the output r-th projection subspace:
[0058] Finally, the scores of all randomly projected subspaces are fused to obtain the overall discriminant score of category c:
[0059] Among them, S c Represents the overall discriminant score of category c. The model compares the discriminant scores of all categories and selects the category with the highest score as the final classification result.
[0060] In order to enable the attention module to fully integrate image features and text features to form multimodal features, the joint optimization of multiple supervisory signals is realized during the training process, involving the following total loss function L total :
[0061] in, is the weight parameter, .
[0062] In order to enhance the alignment ability of image and text features in the feature space, the image-text contrast loss is introduced, and the InfoNCE loss function is used to implement the contrast loss:
[0063] Among them, L contrast represents the contrast loss function; v i Represents image features; t i and t k represents text features; B represents the number of training samples; sim(,) represents cosine similarity; Represents the temperature coefficient.
[0064] In order to maintain the representation consistency between the image features before guiding the attention module and the reference image (REF) features, a distillation constraint is introduced:
[0065] Among them, L distill represents the distillation constraint function; represents the reference image features; represents the two-norm.
[0066] After fusing the image and text features, the prediction score is obtained through random projection and score prediction , calculate the cross entropy loss between the predicted score and the true score:
[0067] Among them, L cross represents the cross entropy loss function; y i,c represents the true score; Represents the prediction score.
[0068] The joint optimization mechanism of the total loss function in this scheme aims to improve the generalization ability, stability and semantic alignment of the classification model that integrates multimodal features, and reduce the forgetting problem in continuous learning through distillation.
[0069] The image to be recognized is input into the trained classification model, which undergoes image encoding, text encoding, guided attention fusion, random projection and Gram matrix classification score in sequence, and finally outputs the category recognition result of the image.
[0070] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An open vocabulary multi-task image classification method based on continuous learning, characterized by: The following steps are involved: Step 1: Obtain original image data, pre-process the original image data to obtain corresponding text information, input the text information into a text encoder, and obtain text features; Step 2: Input the original image data into the image encoder to obtain image features; Step 3: Input the text features and image features into the guided attention module, perform weighted integration on the image features, and obtain a multimodal feature representation that is semantically consistent with the text features; Step 4: Input the multimodal features into the random projection module for random projection, and obtain the activated features through the nonlinear activation function; The prediction module takes the mean of the activation features of each category to generate a class prototype vector, and the nonlinear activation function is input into the Gram matrix to obtain the category of the original image data.
2. The open vocabulary multi-task image classification method based on continuous learning according to claim 1, characterized in that The step 1 specifically includes the following steps: Obtain raw image data from a variety of public datasets. Each public dataset represents a classification task, and build a predefined template for each classification task. The text information of the original image data is extracted through a predefined module. The text information is first segmented and converted into a token sequence of length T. It is then mapped into an embedding vector sequence, and then a learnable positional encoding is added to form the input sequence: Among them, W 0 represents the input sequence; w t represents the embedding vector sequence of the tth word; p t represents the positional encoding corresponding to the t-th embedding vector sequence; t=1,2,...,T; Input sequence W 0 The text encoder is then input for context modeling. The text encoder outputs a special token [EOS] at the end of the sentence as a text feature that aggregates the text information of the entire sentence: Among them, f text Represents text features; TransformerEncoder represents the operation of the text encoder; ( ) [EOS] Indicates a special token at the end of a sentence.
3. The open vocabulary multi-task image classification method based on continuous learning according to claim 1, characterized in that The image encoder in step 2 includes L stacked Transformer encoding blocks, wherein the first part of the Transformer encoding blocks is used as a frozen image encoder, and the second part of the Transformer encoding blocks is used as a fine-tuned image encoder; The step 2 specifically includes the following steps: The original image data I is divided into image blocks of size P×P, and the total image blocks, flatten each image block into a vector, forming a sequence {x1,x2,...,x N }, where each , i=1,2,...,N; H, W, C are the height, width, and number of channels of the original image data I, respectively. is the set of real numbers; Each image block is projected through a linear mapping layer and mapped to a unified feature dimension space , D is the feature dimension, and the embedding vector corresponding to the image block is obtained: in, is the embedding vector corresponding to the i-th image block; is the linear projection matrix; is the two-dimensional position encoding vector corresponding to the i-th image block; The embedding vector sequence of N image patches It is first fed into the first Transformer encoder block, and each Transformer encoder block processes: Among them, Z l represents the output of the lth Transformer encoding block; Z l-1 represents the output of the l-1th Transformer encoding block; FFN represents the operation of the feedforward network; MHSA represents the operation of the multi-head self-attention module; l=1,2,...,L; In the L stacked Transformer encoding blocks, the original image data I first passes through the frozen image encoder to extract intermediate features: in, is the intermediate feature; x I is the original image data; For frozen image encoder; Next, the intermediate feature z int Input the fine-tuned image encoder to obtain the final image features: Among them, v is the image feature; for fine-tuning the image encoder; Represents the fine-tuning parameters.
4. The open vocabulary multi-task image classification method based on continuous learning according to claim 1, characterized in that The step 3 specifically includes the following steps: The text feature f text Mapped to query vector Q, the image feature v is mapped to key vector K and value vector V: in, is a learnable parameter matrix; Flatten() represents the flattening of the architecture image features into a two-dimensional tensor of shape B×(H×W)×D, where B is the batch size, D is the feature dimension, and H and W are the height and width of the original image data, respectively. Take the dot product of the query vector Q and the key vector K to get the attention score S: Where d is the dimension of the query vector Q and the key vector K; Perform Softmax normalization on the attention score S to obtain the attention weight A: Use the attention weight A to perform weighted summation on the value vector V to obtain the fused multimodal feature F attn : After removing redundant dimensions, the final multimodal features are obtained: Among them, f attn is the final multimodal feature; Squeeze( ) represents the operation of removing redundant dimensions.
5. The open vocabulary multi-task image classification method based on continuous learning according to claim 4, characterized in that In step 4, the step of inputting the multimodal features into the random projection module for random projection and obtaining the activated features through the nonlinear activation function includes: Use R linearly invariant random projection matrices For multimodal features f attn To map: Among them, z r is the activation feature under the r-th projection subspace; P r represents the rth projection matrix, R is the number of projection matrices; is a nonlinear activation function; M represents the projection dimension.
6. The open vocabulary multi-task image classification method based on continuous learning according to claim 5, characterized in that In step 4, the step of taking the mean of the activation features of each category through the prediction module to generate a class prototype vector, inputting the nonlinear activation function into the Gram matrix, and obtaining the category of the original image data includes: For category c, the prediction module collects z r Activation feature representation in training samples , and take the average value to form the class prototype vector: Among them, p c,r is the class prototype vector of the c-th category; is the i-th activated feature in category c under the r-th projection subspace, i=1,2,...,N c , N c is the total number of features in category c; The activation feature z is calculated by r With each class prototype vector p c,r The inner product similarity between them is used to construct the Gram matrix to obtain the score of the output r-th projection subspace: The scores of all randomly projected subspaces are combined to obtain the overall discriminant score of category c: Among them, S c represents the overall discriminant score for category c.
Citation Information
Patent Citations
Open vocabulary target detection method, system and device and storage medium
CN118230329A
Cross-modal multi-scale information fusion method and system supporting design knowledge pushing
CN120408491A
Training image processing neural networks using cross-modal alignment
WO2025104314A1