Class increment image classification method and system based on multi-modal pre-training model

By adopting a multimodal pre-trained model approach in class-incremental image classification, and using task-specific adapters and hybrid mapping modules for fine-tuning and feature calibration, the problem of catastrophic forgetting in existing technologies is solved, and cross-task category confusion is alleviated and output feature accuracy is improved.

CN120673143APending Publication Date: 2025-09-19SUN YAT SEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510760095.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies suffer from the catastrophic forgetting problem in class-incremental image classification, making it difficult to effectively adapt to new categories or changes in data distribution, resulting in the model's performance improving on new data while its performance on old data drops sharply.

Method used

This paper adopts a class-incremental image classification method based on a multimodal pre-trained model, achieving continuous learning through inter-task parameter isolation and a two-stage training approach. The specific steps include a within-task training phase and a cross-task feature calibration phase. The pre-trained vision-language model is fine-tuned using task-specific adapters, and cross-entropy loss training is performed through a hybrid mapping module and a text encoder. Finally, classification is performed through an uncertainty-guided reasoning strategy.

Benefits of technology

It effectively avoids the interference of new task knowledge on old task knowledge, alleviates the problem of cross-task category confusion, improves the accuracy of selecting output features, and ensures balanced performance improvement of the model on new and old data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673143A_ABST
    Figure CN120673143A_ABST
Patent Text Reader

Abstract

The invention discloses a class increment image classification method and system based on a multi-modal pre-training model, and the method comprises the steps: obtaining a class increment learning data set, and for each task, model training comprises two stages: an intra-task training stage and a cross-task fine tuning stage; in-task training stage: for the data of the current task, performing fine adjustment on the pre-trained visual language model by adopting a task-specific adapter to realize classification among categories in the task; a cross-task fine tuning stage: introducing a mapping module for image feature expression, mapping features of a specific feature space of a task to a feature space shared by the task, and realizing cross-task category separability; during reasoning, a reasoning strategy based on prediction uncertainty is adopted for image classification. According to the method, the problem of category confusion existing across tasks can be solved, and the accuracy of selecting the output features is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image classification, and specifically relates to a class-incremental image classification method and system based on a multimodal pre-training model. Background Art

[0002] Deep learning models have made significant progress in computer image classification. However, traditional image classification models are typically trained on static datasets and struggle to adapt effectively to new categories or changes in data distribution after training. This leads to a problem known as "catastrophic forgetting," whereby performance on old data decreases dramatically while improving on new data. In real-world applications, data is often continuously generated and may include new, unseen categories. To address this need for continuous learning, class-incremental learning (CIL) has become an important research direction. CIL aims to enable models to gradually learn new categories without forgetting previously learned categories. With the rise of multimodal pre-trained models, leveraging information from other modalities to assist in continuous learning of images has gained increasing attention and demonstrated superior performance. A growing number of methods are exploring the use of visual language models (such as CLIP) for continuous learning due to their robustness and ability to effectively combat forgetting.

[0003] Various methods have been proposed in the prior art. Some methods focus on data replay (Replay-based), that is, while learning new categories, retain a portion of the data of old categories and reuse this old data when training new models. However, this method requires storing old data, which may involve issues such as privacy and storage costs. Other methods focus on regularization (Regularization-based), by introducing additional constraints or penalty terms to limit the model from forgetting old knowledge when learning new knowledge, but this may cause the model to be unable to fully learn new knowledge. Another type of method attempts to dynamically expand the model structure (Dynamic Network Expansion), allocating new model capacity to new categories to reduce interference between knowledge of different tasks. However, this approach faces the problem of selecting the output of specific modules or the fusion of parameters during inference. Summary of the Invention

[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a class-incremental image classification method and system based on a multimodal pre-training model, which realizes continuous learning through parameter isolation between tasks, and a two-stage training method for model learning at the intra-task and cross-task levels respectively, which can solve the problem of category confusion between tasks and improve the accuracy of selecting output features.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a class-incremental image classification method based on a multimodal pre-trained model, comprising the following steps:

[0007] Each task includes an intra-task training phase and a cross-task feature calibration phase.

[0008] In-task training phase:

[0009] Get training images for the task;

[0010] Using a task-specific adapter, fine-tuning a pre-trained visual language model includes an image encoder and a text encoder, inputting a training image into the image encoder to obtain image features, inputting the image features into a hybrid mapping module, and inputting the category text into the text encoder, training the hybrid mapping module and the text encoder with cross-entropy loss, obtaining a trained task-specific adapter, and preliminarily trained hybrid mapping module, freezing the image encoder and text encoder;

[0011] Cross-task feature calibration stage:

[0012] The trained image encoder outputs image features of each category to construct a pseudo-feature distribution. Then, pseudo-features of each category are sampled from the distribution to train a hybrid mapping module. The hybrid mapping module learns different mapping combinations from pseudo-features of multiple categories and aligns the image features of each category with the text features of the corresponding category.

[0013] After completing all tasks, a trained hybrid mapping module is obtained;

[0014] During inference, an inference strategy based on prediction uncertainty is used to classify the input image.

[0015] As a preferred technical solution, the pre-trained visual language model is fine-tuned using a task-specific adapter. Specifically, the task-specific adapter is connected in parallel with the feedforward layer of each layer of the image encoder, as shown in the following formula:

[0016] x o =x i +MLP(x i )+s·W up ·σ(Wdown ·x i )

[0017] Among them, x i is the input vector, i.e. the output of the previous self-attention layer, x o is the output vector, s is the scaling factor, σ is the activation function, is the upsampling layer, is the upsampling layer.

[0018] As a preferred technical solution, the hybrid mapping module includes a gating network and multiple mapping modules, the gating network includes a linear layer and a Softmax activation function σ, and the mapping modules are multiple multi-layer induction machines.

[0019] As a preferred technical solution, the step of inputting the image features into the hybrid mapping module and inputting the category text into the text encoder includes:

[0020] For an input image feature z, it is obtained after passing through the gating network:

[0021] g m =σ(W g z)

[0022] Among them, W g is the weight of the linear layer, g m Represents the weight score assigned to each mapping module by the gating network for the input feature;

[0023] The mapping result is obtained by weighted summation based on the weight score:

[0024]

[0025] Among them, P m (·) is the mapping module, and m is the number of multi-layer sensing machines.

[0026] As a preferred technical solution, the hybrid mapping module and the text encoder are trained with cross entropy loss, as shown in the following formula:

[0027]

[0028] in, Represents the category text features encoded by the text encoder as the classification head, θ A ,θ MoP Represents the parameters of the adapter and hybrid mapping modules respectively.

[0029] As a preferred technical solution, the method of constructing a pseudo feature distribution using the image features of each category output by the trained image encoder and sampling the pseudo features of each category includes:

[0030] The mean and covariance matrix of the image feature vectors of the training samples of each class are calculated and saved, and the characteristic Gaussian distribution of each class is constructed. The Gaussian distribution is used as a pseudo feature distribution and sampled from it to obtain pseudo features; the pseudo features of each category are used to train the hybrid mapping module to calibrate and fine-tune the category feature expressions originally belonging to the feature space of each task to the same cross-task feature space.

[0031] As a preferred technical solution, the reasoning strategy based on prediction uncertainty includes:

[0032] For a test image belonging to task t, the test image is passed through an image encoder to obtain pre-calibrated image features, and the pre-calibrated image features are passed through a hybrid mapping module to obtain calibrated image features. The image encoder is fine-tuned by a task-specific adapter;

[0033] Calculate the probability distribution of the calibrated image features and all categories, calculate the uncertainty based on the probability distribution, and obtain a first value; calculate the uncertainty of the image features of the test image before calibration and the probability distribution of all categories to obtain a second value, and calculate the uncertainty difference of the corresponding image features after the hybrid mapping module, that is, the difference between the first value and the second value;

[0034] After completing all tasks, the first value and the difference are used to perform image classification according to the feature selection strategy.

[0035] As a preferred technical solution, the calculated calibrated image features and the probability distribution of all categories are calculated, and the uncertainty is calculated according to the probability distribution to obtain the first value, as shown in the following formula:

[0036]

[0037] Where p=[p1,p2,…,p C ] represents the probability distribution of the calibrated image features over all C categories.

[0038] As a preferred technical solution, the image classification is performed according to the feature selection strategy, as shown in the following formula:

[0039]

[0040] in, represents the indicator function, m is a constant greater than 0, ε t (p) is the first value, that is, the entropy value of the probability distribution of image features and text features of all categories, Δε t is the difference between the first value and the second value;

[0041] Based on selected features Calculate the similarity with the text features of all categories and select the one with the highest similarity as the classification result.

[0042] In a second aspect, the present invention further provides a class-incremental image classification system based on a multimodal pre-training model, which is applied to the class-incremental image classification method based on a multimodal pre-training model, and includes a first execution module and a second execution module;

[0043] The first execution module is used to execute each task, each task includes an intra-task training phase and a cross-task feature calibration phase,

[0044] In-task training phase:

[0045] Get training images for the task;

[0046] Using a task-specific adapter, fine-tuning a pre-trained visual language model includes an image encoder and a text encoder, inputting a training image into the image encoder to obtain image features, inputting the image features into a hybrid mapping module, and inputting the category text into the text encoder, training the hybrid mapping module and the text encoder with cross-entropy loss, obtaining a trained task-specific adapter, and preliminarily trained hybrid mapping module, freezing the image encoder and text encoder;

[0047] Cross-task feature calibration stage:

[0048] The hybrid mapping module learns different weight distributions from pseudo features of multiple categories and aligns the image features of each category with the text features of the corresponding category;

[0049] After completing all tasks, a trained hybrid mapping module is obtained;

[0050] The second execution module is used to perform reasoning. During reasoning, an inference strategy based on prediction uncertainty is used to classify the input image.

[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0052] (1) The present invention adopts a network structure that combines a hybrid mapping module with a pre-trained visual language model. This can avoid the interference of new task knowledge on old task knowledge during the continuous learning process. At the same time, the two-stage training method is implemented in a single task to learn the model at the intra-task and cross-task levels, which can alleviate the cross-task category confusion problem.

[0053] (2) The present invention uses an uncertainty-guided reasoning strategy to design a strategy that selects image features with smaller entropy values ​​and a decreasing entropy value to select image features for final classification. Compared with other continuous learning methods with model structure expansion, it can select output features more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0055] Figure 1 Flowchart of a class-incremental image classification method based on a multimodal pre-training model according to an embodiment of the present invention;

[0056] Figure 2 It is a diagram of the reasoning strategy of an embodiment of the present invention;

[0057] Figure 3 This is a graphical representation of the uncertainty law based on the inference strategy of an embodiment of the present invention on the ImageNet-R dataset;

[0058] Figure 4 Schematic diagram of the structure of a class-incremental image classification system based on a multimodal pre-training model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0060] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0061] See also Figure 1 , this embodiment provides a class-incremental image classification method based on a multimodal pre-training model:

[0062] There are T tasks in represents image data, Represents the image category label. Note that the data categories of each task are disjoint. The category set of the data from task 1 to t is hereinafter referred to as The goal of incremental learning is to ensure that a learning system f(x) can continuously learn new tasks without forgetting the knowledge learned from previous tasks, even if T task data are not available simultaneously. This means that the system can still distinguish between old class data. This example focuses on the setting where no old class samples are replayed.

[0063] In this embodiment, the visual language model CLIP based on image-text matching serves as the backbone of the pre-trained model. CLIP consists of two parts: an image encoder f(x) and a text encoder h(w). f(x) encodes the input image into a D-dimensional feature vector, while h(w) similarly encodes the input text into a D-dimensional text feature vector. During classification, the image category probability distribution is calculated using the following formula:

[0064]

[0065] Where z = f(x), cos() represents the calculation of cosine similarity, τ is the learnable temperature coefficient, w c is the word embedding of the input text of category c (using the template “a photo of a[class name]”).

[0066] For each task, model training includes a within-task training phase and a cross-task feature calibration phase.

[0067] First, the task-specific training phase. First, obtain the training images for task t and insert the task-specific adapter into the feed-forward layer of each layer of the image encoder in parallel with it:

[0068] x o =x i +MLP(x i )+s·W up ·σ(W down ·x i )

[0069] Among them, x i is the input vector, i.e. the output of the previous self-attention layer, x o is the output vector, s is the scaling factor, σ is the activation function, is the upsampling layer, is the upsampling layer.

[0070] It is worth explaining that a task-specific lightweight adapter module is connected in parallel next to the feed-forward layer (FFN) of each pre-trained, fixed-parameter image encoder. These adapter modules focus on learning unique feature representations for the current task, and their parameters are independent across different tasks, thus achieving parameter isolation between tasks and avoiding catastrophic forgetting.

[0071] In this training phase, the task-shared hybrid mapping module is connected to the image encoder for training together. The purpose is to let it learn the distribution of real data, obtain a better weight initialization point, and be more conducive to cross-task feature calibration training. The image features output by the encoder are then input into the hybrid mapping module (MoP), which consists of two parts: a gating network and multiple mapping modules. The gating network consists of a linear layer and Softmax activation function σ; mapping module {P m |m=1,…,M} is M MLP. For an input feature z, after passing through the gating network, we get:

[0072] g=σ(W g z)

[0073] where g=[g1,g2,…,g M ] represents the weight score assigned to each mapping module by the gating network for the input feature. Based on this score, the mapping result can be obtained by weighted summation:

[0074]

[0075] After completing the above steps, training is performed by minimizing the cross entropy loss through back propagation, as follows:

[0076]

[0077] in, Represents the category text features encoded by the text encoder as the classification head, θ A ,θ MoP Represents the parameters of the adapter and hybrid mapping modules respectively.

[0078] In this implementation, after completing the in-task training phase of training, the trained task-specific adapter and the preliminarily trained hybrid mapping module are obtained, and then the pre-trained image encoder and text encoder are frozen, and only the preliminarily trained hybrid mapping module is trained.

[0079] Second, the cross-task feature calibration phase. The features of the image encoder output obtained by the in-task training phase are used to construct the characteristic Gaussian distribution of each class. As the pseudo feature distribution of the category. Specifically, the mean and covariance matrix of the feature vectors of the training samples of each class are calculated and saved. During the subsequent cross-task feature calibration phase of each task, pseudo features are sampled from the pseudo feature distribution of all classes and used together to train the hybrid mapping module, so that it can calibrate and fine-tune the category feature expressions that originally belonged to the feature space of each task to the same cross-task feature space, thereby achieving cross-task category separability. The hybrid mapping module learns the differences in features of different classes, so that the gating network can learn different weight distributions according to the features of different categories, and then use different hybrid methods for mapping, which has a better calibration effect.

[0080] This involves a pseudo-feature replay step, where the sample feature vectors learned by the task-specific adapter can be used to calculate the pseudo-Gaussian distribution of the features of each class by calculating the mean and covariance matrix. This distribution can be saved for cross-task representation calibration.

[0081] Figure 2 Shown is a diagram of model reasoning in an embodiment.

[0082] Inference phase. After training T tasks, we get T sets of task-specific adapters. Then, for an input image, we can get T calibrated image features {z ′ t |t=1,…,T}. During inference, the most appropriate one needs to be selected from the T image features. That is, the feature corresponding to the task to which the input image belongs is used to calculate the similarity with the category text for classification.

[0083] During reasoning in this embodiment, an inference strategy based on prediction uncertainty is used to classify the input image. After learning T tasks, the uncertainty-guided inference strategy can obtain T groups of task-specific adapters. Then, for a test image, T image features will be obtained. The strategy will select the most appropriate feature from them to calculate the similarity with the category text features for classification. The uncertainty of the similarity calculated between the image features calibrated in the hybrid mapping module and the text features of all categories will be lower, which is specifically manifested as a lower entropy value. In addition, for a test image belonging to task t, compared with the entropy value corresponding to the feature before calibration, the features output by task t will usually show a decrease in entropy after calibration, while the features output by non-task t are more inclined to show an increase in entropy. Based on this rule, a strategy is designed to select image features with smaller entropy values ​​and a decrease in entropy values ​​to select image features for final classification.

[0084] To more clearly illustrate this embodiment, the uncertainty-guided reasoning strategy can be divided into the following sections:

[0085] (1) Select features based on the magnitude of the prediction uncertainty. Specifically, for a test image belonging to task t, the test image is passed through an image encoder fine-tuned by a task-specific adapter to obtain pre-calibrated image features, the pre-calibrated image features are passed through a hybrid mapping module to obtain calibrated image features, and then the probability distribution of the calibrated image features and all categories is calculated. The uncertainty is calculated based on the probability distribution to obtain a first value. In this embodiment, the uncertainty is expressed as the entropy value of the probability distribution:

[0086]

[0087] Where p=[p1,p2,…,p C ] represents the probability distribution of the calibrated image features over all C categories.

[0088] Prediction uncertainty is the uncertainty of the output probability distribution of an image across all categories. Higher uncertainty indicates that the model is less confident in its judgment of the sample, while lower uncertainty indicates greater confidence. Therefore, image features corresponding to probability distributions with lower uncertainty are more likely to come from the task corresponding to the input image.

[0089] (2) Assist in selecting features based on the change of uncertainty. Specifically, for the image features {z t |t=1,…,T}, similarly, the entropy of the probability distribution is used to calculate the entropy of the probability distribution of all categories ε t (q), calculate the uncertainty difference of the image features before and after calibration, that is, Δε t =ε t (p)-ε t (q).

[0090] After completing all T tasks, use the ε obtained from all tasks t (p) and Δε t Feature selection classification is performed according to the feature selection strategy. Specifically, Figure 2 As shown, there are two arrows pointing to the “feature selection function”, one is the uncertainty difference Δε t , one is the entropy value ε of the probability distribution of the calibrated feature t (p), and then select the task with the minimum calculation result in this batch of tasks through feature selection strategy. As the prediction taskid of the test image.

[0091] After practice, if the input test image belongs to task i, then Δε i is generally less than 0, that is, the uncertainty is reduced; and for j≠i, Δε j It tends to be greater than 0, which means that the uncertainty increases. Figure 3The figure shows the performance of this phenomenon under the task setting of the dataset ImageNet-R 10. The darker the color, the greater the proportion of samples with reduced entropy in the task. Obviously, the color of the diagonal position is darker than that of other positions. Based on this, the feature selection strategy of this embodiment is as follows: Figure 2 The selection feature function in is designed as:

[0092]

[0093] in, Represents the indicator function, and m is a constant greater than 0. For a calibrated image feature that does not match the task of the input image, its Δε is more likely to be greater than 0, so that even if its original ε t (p) unexpectedly appears smaller and leads to misjudgment, then adding a positive value can correct this problem. Finally, based on the selected features Calculate the similarity with the text features of all categories and select the one with the highest similarity as the classification result.

[0094] The specific parameters of this embodiment are as follows: the pre-trained visual language model uses Open AI CLIP ViT-B-16 version, and the model is trained using the AdamW optimizer. The mapping module in the hybrid mapping module uses MLP, and M is set to 3. In the second stage of training, 256 pseudo features are sampled from the pseudo feature Gaussian distribution of each class. Experimental verification shows that m is set to 10 -2 The order of magnitude is relatively appropriate, and setting it to 0.02 has a better effect.

[0095] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.

[0096] Based on the same idea as the multimodal pre-trained model-based incremental image classification method in the above-mentioned embodiment, the present invention also provides a multimodal pre-trained model-based incremental image classification system, which can be used to execute the above-mentioned multimodal pre-trained model-based incremental image classification method. For ease of explanation, the structural diagram of the embodiment of the multimodal pre-trained model-based incremental image classification system only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0097] See also Figure 4, in another embodiment of the present application, a class-incremental image classification system 10 based on a multimodal pre-trained model is provided, the system comprising a first execution module 11 and a second execution module 12;

[0098] The first execution module 11 is used to execute each task, each task includes an intra-task training phase and a cross-task feature calibration phase,

[0099] In-task training phase:

[0100] Get training images for the task;

[0101] Using a task-specific adapter, fine-tuning a pre-trained visual language model includes an image encoder and a text encoder. The training image is input into the image encoder to obtain image features. The image features are input into a hybrid mapping module, and the category text is input into the text encoder. The hybrid mapping module and the text encoder are trained with cross-entropy loss to obtain a trained task-specific adapter and a preliminarily trained hybrid mapping module, and the image encoder and text encoder are frozen.

[0102] Cross-task feature calibration stage:

[0103] The trained image encoder outputs image features of each category and performs feature sampling to construct pseudo features of the corresponding category. The hybrid mapping module learns different weight distributions from pseudo features of multiple categories and aligns the image features of each category with the text features of the corresponding category.

[0104] After completing all tasks, a trained hybrid mapping module is obtained;

[0105] The second execution module 12 is used to perform reasoning. During reasoning, the trained task-specific adapter classifies the input image using a reasoning strategy based on prediction uncertainty.

[0106] It should be noted that the class-incremental image classification system based on a multimodal pre-training model of the present invention corresponds one-to-one to the class-incremental image classification method based on a multimodal pre-training model of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the class-incremental image classification method based on a multimodal pre-training model are all applicable to the embodiment of the class-incremental image classification method based on a multimodal pre-training model. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.

[0107] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0108] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A class-incremental image classification method based on a multimodal pre-trained model, characterized in that: include: Each task includes an intra-task training phase and a cross-task feature calibration phase. In-task training phase: Get training images for the task; Using a task-specific adapter, fine-tuning a pre-trained visual language model includes an image encoder and a text encoder, inputting a training image into the image encoder to obtain image features, inputting the image features into a hybrid mapping module, and inputting the category text into the text encoder, training the hybrid mapping module and the text encoder with cross-entropy loss, obtaining a trained task-specific adapter, and preliminarily trained hybrid mapping module, freezing the image encoder and text encoder; Cross-task feature calibration stage: The trained image encoder outputs image features of each category to construct a pseudo-feature distribution. Then, pseudo-features of each category are sampled from the distribution to train a hybrid mapping module. The hybrid mapping module learns different mapping combinations from pseudo-features of multiple categories and aligns the image features of each category with the text features of the corresponding category. After completing all tasks, a trained hybrid mapping module is obtained; During inference, an inference strategy based on prediction uncertainty is used to classify the input image.

2. The class-incremental image classification method based on a multimodal pre-training model according to claim 1, characterized in that: The pre-trained visual language model is fine-tuned using the task-specific adapter. Specifically, the task-specific adapter is connected in parallel with the feed-forward layer of each layer of the image encoder, as shown in the following formula: x o =x i +MLP(x i )+s·W up ·σ(W down ·x i ) Among them, x i is the input vector, i.e. the output of the previous self-attention layer, x o is the output vector, s is the scaling factor, σ is the activation function, is the upsampling layer, is the upsampling layer.

3. The class-incremental image classification method based on a multimodal pre-training model according to claim 1, characterized in that: The hybrid mapping module includes a gating network and multiple mapping modules, the gating network includes a linear layer and a Softmax activation function σ, and the mapping modules are multiple multi-layer induction machines.

4. The class-incremental image classification method based on a multimodal pre-training model according to claim 3, characterized in that: The step of inputting the image features into the hybrid mapping module and inputting the category text into the text encoder includes: For an input image feature z, it is obtained after passing through the gating network: g m =σ(W g With) Among them, W g is the weight of the linear layer, g m Represents the weight score assigned to each mapping module by the gating network for the input feature; The mapping result is obtained by weighted summation based on the weight score: Among them, P m (·) is the mapping module, and m is the number of multi-layer sensing machines.

5. The class-incremental image classification method based on a multimodal pre-training model according to claim 1, characterized in that: The hybrid mapping module and the text encoder are trained with cross entropy loss as follows: in, Represents the category text features encoded by the text encoder as the classification head, θ A ,θ MoP Represents the parameters of the adapter and hybrid mapping modules respectively.

6. The class-incremental image classification method based on a multimodal pre-training model according to claim 1, characterized in that: The method of constructing a pseudo feature distribution by using the image features of each category output by the trained image encoder and sampling the pseudo features of each category includes: The mean and covariance matrix of the image feature vectors of the training samples of each class are calculated and saved, and the characteristic Gaussian distribution of each class is constructed. The Gaussian distribution is used as a pseudo feature distribution and sampled from it to obtain pseudo features; the pseudo features of each category are used to train the hybrid mapping module to calibrate and fine-tune the category feature expressions originally belonging to the feature space of each task to the same cross-task feature space.

7. The class-incremental image classification method based on a multimodal pre-training model according to claim 1, characterized in that: The reasoning strategy based on prediction uncertainty includes: For a test image belonging to task t, the test image is passed through an image encoder to obtain pre-calibrated image features, and the pre-calibrated image features are passed through a hybrid mapping module to obtain calibrated image features. The image encoder is fine-tuned by a task-specific adapter; Calculate the probability distribution of the calibrated image features and all categories, calculate the uncertainty based on the probability distribution, and obtain a first value; calculate the uncertainty of the image features of the test image before calibration and the probability distribution of all categories to obtain a second value, and calculate the uncertainty difference of the corresponding image features after the hybrid mapping module, that is, the difference between the first value and the second value; After completing all tasks, the first value and the difference are used to perform image classification according to the feature selection strategy.

8. The class-incremental image classification method based on a multimodal pre-training model according to claim 7, characterized in that: The calculated calibrated image features and probability distribution of all categories are calculated, and uncertainty is calculated according to the probability distribution to obtain a first value, as shown in the following formula: Where p=[p1,p2,…,p C ] represents the probability distribution of the calibrated image features over all C categories.

9. The class-incremental image classification method based on a multimodal pre-training model according to claim 7, characterized in that: The image classification is performed according to the feature selection strategy as follows: in, Represents the indicator function, m is a constant greater than 0, ε t (p) is the first value, that is, the entropy value of the probability distribution of image features and text features of all categories, Δε t is the difference between the first value and the second value; Based on selected features Calculate the similarity with the text features of all categories and select the one with the highest similarity as the classification result.

10. A class-incremental image classification system based on a multimodal pre-trained model, characterized in that: The class incremental image classification method based on the multimodal pre-training model applied to any one of claims 1 to 9 comprises a first execution module and a second execution module; The first execution module is used to execute each task, each task includes an intra-task training phase and a cross-task feature calibration phase, In-task training phase: Get training images for the task; Using a task-specific adapter, fine-tuning a pre-trained visual language model includes an image encoder and a text encoder, inputting a training image into the image encoder to obtain image features, inputting the image features into a hybrid mapping module, and inputting the category text into the text encoder, training the hybrid mapping module and the text encoder with cross-entropy loss, obtaining a trained task-specific adapter, and preliminarily trained hybrid mapping module, freezing the image encoder and text encoder; Cross-task feature calibration stage: The hybrid mapping module learns different weight distributions from pseudo features of multiple categories and aligns the image features of each category with the text features of the corresponding category; After completing all tasks, a trained hybrid mapping module is obtained; The second execution module is used to perform reasoning. During reasoning, an inference strategy based on prediction uncertainty is used to classify the input image.

Citation Information

Cited By

  • Target detection method and device, storage medium and computer equipment

    CN121564303A

  • Target detection method and device, storage medium and computer device

    CN121564303B