Feature Uniform Alignment Method for Pre-trained Multimodal Models Based on Meta-Learning
By introducing a dual prompt framework and alignment and uniformity constraints in the multimodal model, combined with meta-learning and weight adaptive adjustment, the problem of insufficient alignment and uniformity of feature distribution in the fine-tuning of the multimodal model is solved, and the generalization ability and performance of the model are significantly improved.
Patent Information
- Application Number
- CN202411510421.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing multimodal model fine-tuning method based on prompt learning has problems such as insufficient alignment and uniformity of the internal feature distribution of single mode, insufficient task generalization ability, and inapplicable regularization constraints.
A pre-trained multimodal model feature uniform alignment method is adopted based on meta-learning. By introducing a dual prompt framework and alignment and uniformity constraints, combined with the weight adaptive adjustment module, a meta-enhanced contrast learning model is built to optimize the prompt parameters and weight adaptive adjustment model parameters.
Improve modal alignment and feature uniformity, enhance the generalization ability of the model, reduce the risk of overfitting, and significantly improve the performance of multimodal models in downstream tasks.
Smart Images

Figure CN119360112B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence. More specifically, it relates to a method for uniformly aligning features of a pre-trained multi-modal model based on meta-learning. Background Art
[0002] In recent years, in the field of artificial intelligence, researchers have been continuously exploring how to better utilize the connections between multi-modal data to achieve a more intelligent information processing model. Among them, vision-language pre-trained models (such as CLIP and ALIGN) have received extensive attention and rapid development. With the help of modern computer storage technology and GPU computing power, these models are trained on a large amount of image-text pair data, showing powerful zero-shot inference capabilities and achieving remarkable results in tasks such as image classification, object detection, and segmentation. With the success of these pre-trained models, how to efficiently fine-tune them to adapt to specific downstream tasks has become the focus of research.
[0003] Existing fine-tuning methods are mainly divided into three categories: adapters, prompt learning, and low-rank adaptation (LoRA). Among them, the LoRA technique improves the efficiency of model fine-tuning by decomposing the weight matrix into the product of low-rank matrices and fine-tuning these low-rank parameter matrices. However, the low-rank decomposition assumption of LoRA does not always apply to all tasks, especially when dealing with very complex or datasets with strong non-linear relationships. The adapter method adds learnable additional layers on specific layers of the pre-trained model to adapt to different task requirements. However, this method sometimes introduces additional model complexity, thus increasing the computational overhead. Prompt learning draws on the prompting technique in the field of natural language processing and modifies the input of the model so that it takes into account the specific requirements of the task at the input stage. Currently, prompt learning has become the mainstream method for fine-tuning multi-modal large models. However, there are some key problems in the existing prompt learning-based methods:
[0004] (1) Insufficient alignment and uniformity of the feature distribution within a single modality: Current prompt learning methods mainly focus on using scarce labeled domain data to design various efficient prompt forms to achieve feature space alignment between the visual and language modalities. However, these methods often ignore the feature distribution within a single modality. When the embeddings of each modality are more uniformly arranged in the latent space, the alignment between modalities will be more effective.
[0005] (2)Insufficient task generalization ability: Since soft prompts tend to prioritize knowledge of specific tasks, the model is prone to overfitting on the target task, resulting in a decline in task generalization ability. Although some methods alleviate the overfitting problem in multimodal fine-tuning by introducing new regularization constraints, the regularization constraints are not effective for all tasks. How to adaptively adjust the strength of the regularization term according to the characteristics of different samples remains an urgent problem to be solved. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for uniformly aligning features of a pre-trained multimodal model based on meta-learning. By introducing a dual prompt framework and alignment and uniformity constraints, the feature extraction ability for images and texts is improved, thereby significantly enhancing the performance of the multimodal model in downstream tasks.
[0007] To achieve the above invention purpose, the method for uniformly aligning features of a pre-trained multimodal model based on meta-learning of the present invention includes the following steps:
[0008] S1: Construct and pre-train a multimodal model according to actual needs, including a word embedding encoding module, a text encoder block, an embedding encoding module, and an image encoder, where:
[0009] The word embedding encoding module is used to perform word embedding encoding on the input text m to obtain word embeddings and send them to the text encoder;
[0010] The text encoder is used to perform encoding processing on the word embeddings to obtain text embeddings t;
[0011] The block embedding encoding module is used to perform block embedding encoding on the input image x to obtain block embeddings and send them to the image encoder;
[0012] The image encoder is used to perform encoding processing on the block embeddings to obtain image embeddings z;
[0013] S2: Collect M image samples and corresponding texts according to the actual application scenario of the multimodal model to form a training data set S, and then construct a meta data set based on the training data set S
[0014] S3: Taking the pre-trained multimodal model as a benchmark model, replacing the word embedding encoding module with a text-side prompt module, replacing the block embedding encoding module with an image-side prompt module, and adding a weight adaptive adjustment module, thereby constructing a meta-enhanced contrastive learning model, where:
[0015] The text - end prompt module is used to perform word - embedding encoding on the input text m to obtain word embeddings Among them, the parameters of the word - embedding encoding adopt the parameters of the pre - trained word - embedding encoding module. Then, the learnable text prompt vector v = [v 1 ,v 2 ,…,v L and the true class label y are concatenated to obtain the text prompt p = [v 1 ,v 2 ,…,v L ,y]. Then, the text prompt p is concatenated with the word embeddings to obtain the text - prompt embedding m′ and output it to the text encoder;
[0016] The image - end prompt module is used to perform image - embedding encoding on the input image x to obtain patch embeddings Among them, the parameters of the image - embedding encoding adopt the parameters of the pre - trained image - embedding encoding module. Then, the learnable image prompt vector u = [u 1 ,u 2 ,…,u L is inserted into the image - patch class token CLS and the image representation patch E obtained by the image - embedding encoding to obtain the image prompt q = [CLS,u 1 ,u 2 …,u L ,E]. Then, the image prompt q is concatenated with the patch embeddings to obtain the image - prompt embedding x′ and output it to the image encoder;
[0017] The weight adaptive adjustment module is used to receive the text embedding t output by the text encoder and the image embedding z output by the image encoder, and estimate the loss - function weight λ;
[0018] S4: Randomly initialize the prompt parameters w (0) ={v (0) ,u (0)}, the parameters θ (0) of the weight adaptive adjustment model;
[0019] S5: Let the training round t = 1;
[0020] S6: Input the training samples in the training dataset S into the meta - enhanced contrastive learning model, and calculate the loss function using the following formula
[0021]
[0022] where λ i represents the loss - function weight obtained by the weight adaptive adjustment module for the i - th training sample in the training dataset S; Denote the multimodal model loss function of the $i$-th training sample in the training dataset $S$, where $i = 1, 2, \ldots, B$, and $B$ represents the number of samples in the training dataset $S$. The calculation formula is as follows;
[0023]
[0024] where, denotes the probability that the image sample $x$ in the $i$-th training sample i is predicted to belong to class $c$, denotes the predicted class label of the image $x$ i ;
[0025] Denote the alignment loss function of the $i$-th training sample in the training dataset $S$. The calculation formula is as follows:
[0026]
[0027] where, $y$ i denotes the true label of the image sample $x$ in the $i$-th training sample i , $c$ represents the class serial number; $1$ [yi=c] denotes a binary variable, that is, when $y$ i $= c$, $1$ [yi=c] $= 1$, otherwise $1$ [yi=c] $= 0$; $z$ i and $z$ j respectively denote the image embeddings corresponding to the image samples $x$ i and $x$ j in the training dataset $S$ output by the image encoder, where $j = 1, 2, \ldots, B$; $\| \|$ 2 denotes the calculation of the L2 norm;
[0028] Denote the uniformity loss function of the $i$-th training sample in the training dataset $S$. The calculation formula is as follows:
[0029]
[0030] where, $\exp()$ represents the exponential function, $\sigma$ represents the width parameter of the Gaussian kernel function, $t$ i and $t$ j respectively denote the text embeddings corresponding to the texts $m$ i and $m$ j in the training dataset $S$ output by the text encoder;
[0031] Then calculate the gradient of the loss function and update the prompt parameter $w$ using the following formula : (t) :
[0032]
[0033] Among them, α represents a preset learning rate;
[0034] S7: Input the training samples in the meta-dataset into the meta-augmented contrastive learning model, and calculate the loss function using the following formula
[0035]
[0036] where N represents the number of training samples in the meta-dataset , respectively represent the true probability and the predicted probability of the training samples in the meta-dataset on class c, and n = 1, 2, …, N;
[0037] Then calculate the gradient of the loss function and update the parameters θ of the weight adaptive adjustment model using the following formula : (t) :
[0038]
[0039] where β represents a preset learning rate;
[0040] S8: Determine whether the iteration end condition is reached. If so, go to step S9; otherwise, go to step S10;
[0041] S9: Let the training round t = t + 1, and return to step S6;
[0042] S10: Determine the text prompt vector v = [v 1 , v 2 , …, v L and the image prompt vector u = [u 1 , u 2 , …, u L according to the prompt parameters of the current meta-augmented contrastive learning model. Remove the weight adaptive adjustment model from the meta-augmented contrastive learning model to obtain a dual-prompt multi-modal model, and complete the fine-tuning of the multi-modal model.
[0043] The method for uniformly aligning features of a pre-trained multi-modal model based on meta-learning in the present invention uses a pre-trained multi-modal model as a benchmark model, replaces the word embedding encoding module with a text-side prompt module, replaces the patch embedding encoding module with an image-side prompt module, and adds a weight adaptive adjustment model, thereby constructing a meta-augmented contrastive learning model. The prompt parameters, the weight adaptive adjustment model parameters, and the loss function weights are initialized, and then the training data set and the meta-data set are alternately used to learn the prompt parameters and the weight adaptive adjustment model parameters. The text-side prompt module and the image-side prompt module that determine the prompt parameters are added to the multi-modal model to form a dual-prompt multi-modal model, and the fine-tuning of the multi-modal model is completed.
[0044] The present invention has the following beneficial effects:
[0045] 1) Improving modal alignment: By introducing alignment constraints, the present invention reduces the intra-class distance of samples of the same category in the embedding space, enhances the feature alignment effect between the visual and text modalities, and thus improves the performance of the multi-modal model in cross-modal tasks;
[0046] 2) Improving feature uniformity: Through uniformity constraints, the present invention ensures that the model can be evenly distributed in the latent space when learning text features, reduces the risk of over-aggregation of features, improves the generalization ability of the model, especially in complex and variable data sets;
[0047] 3) Adaptive regularization tuning: The present invention introduces an adaptive regularization tuning mechanism based on meta-learning, which can dynamically adjust the strength of the regularization constraint according to the training loss gradient of each sample, effectively avoids the overfitting problem, and significantly improves the adaptability and stability of the model in different tasks;
[0048] 4) Improving the generalization ability of the pre-trained model: By combining the double-layer optimization problem and the online approximation technology, the present invention can efficiently train the weight adaptive adjustment model parameters, enabling the model to quickly adapt when facing different tasks, and thus achieving higher accuracy and stability in multiple downstream tasks. Description of the Drawings
[0049] Figure 1 is a schematic diagram of the technical principle of the present invention;
[0050] Figure 2 is a flowchart of the specific implementation manner of the method for uniformly aligning features of a pre-trained multi-modal model based on meta-learning in the present invention;
[0051] Figure 3 is a structural diagram of the meta-augmented contrastive learning model in the present invention. Specific Embodiment
[0052] The specific embodiments of the present invention will be described below in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0053] To better illustrate the technical solution of the present invention, the fine-tuning idea of the pre-trained multi-modal model of the present invention will be briefly described first.
[0054] Figure 1 It is a schematic diagram of the technical principle of the present invention. As Figure 1 shown, in this embodiment, the multi-modal model is described by taking the prompt learning CLIP model as an example. The traditional CLIP model includes an image feature extraction module, an image encoder, a text feature extraction module, and a text encoder. Given an image-text set M and a sample set X, for each text m i and image x i , the word embedding and the patch embedding are obtained through the text feature extraction module and the image feature extraction module respectively. Then, the text embedding t i and the image embedding z i can be obtained through the text encoder and the image encoder respectively, as follows:
[0055]
[0056] where i represents the training sample serial number, g() represents the text encoding operation, and f() represents the image encoding operation.
[0057] Given C image categories, the prediction probability of an image belonging to each category can be calculated by the softmax function. Therefore, the probability of the image x i belonging to each category is calculated as follows:
[0058]
[0059] where, represents the predicted category label of the image x i . According to the exp() representing the exponential function, the superscript T representing the transpose, τ representing the temperature parameter, and t c represents the text embedding of the c-th category label.
[0060] Define the training loss of the multi-modal model as Usually, the cross-entropy loss is adopted, and its specific calculation formula is as follows:
[0061]
[0062] Among them, B represents the number of samples in the current training data.
[0063] For a multimodal model such as the above-mentioned CLIP model, a prompt framework can be constructed simultaneously at the text end and the image end. The image feature extraction module is replaced by the image-end prompt module, and the text feature extraction module is replaced by the text-end prompt module. That is, image prompts and text prompts are added to the block embedding and word embedding to promote alignment between modalities. Specifically, the text prompt p is represented as a combination of the learnable vector v = [v 1 , v 2 , …, v L and the true class label y, where L represents the length of the learnable vector. Without loss of generality, the length of L in this embodiment is uniformly set to 16. Then, the input of the text encoder, that is, the text prompt p, is:
[0064] p = [v 1 , v 2 , …, v L , y]
[0065] In the training stage, the true class label in each training sample text prompt is the label corresponding to the image in the training sample. In the inference stage, for example, if the multimodal model is used for classification, then the learnable vector can be concatenated with the true class label of each class, and then the text embeddings under different classes are obtained, and the similarity between the image embedding and different text embeddings is calculated to determine the class corresponding to the image. In other multimodal tasks, the true class label in the text prompt can be set according to the actual needs of the task.
[0066] Similarly, since the image X i can also be converted into the form of tokens through a transformer, by inserting the learnable variable u = [u 1 , u 2 , …, u L into the class token CLS of the image patch and the image representation patch E, the input of the image encoder, that is, the image prompt q, is obtained as:
[0067] q = [CLS, u 1 , u 2 …, u L , E]
[0068] Then the text prompt p is concatenated with the word embedding , and the image prompt q is concatenated with the block embedding . The obtained text prompt embedding and image prompt embedding are respectively input into the text encoder and the image encoder for further processing.
[0069] For unified description, the present invention constructs the text - end learnable vector v and the image - end learning variable u into the prompt model parameter w = {v, u}.
[0070] Current prompt learning methods mainly focus on using scarce labeled domain data to design various efficient prompt forms to achieve feature - space alignment between the visual and language modalities. However, these methods often ignore the feature distribution within a single modality. When the feature representations of each modality are arranged more uniformly in the unit hypersphere, the alignment between modalities will be more effective. Inspired by the contrastive learning theory of a single modality, it should be promoted that the modality feature representations learned by the multi - modality large model should satisfy alignment and uniformity. Among them, alignment means that similar samples should be mapped to nearby features. And uniformity means that the feature representations should be roughly evenly distributed on the unit hypersphere to retain as much data information as possible. Starting from the ideas of alignment and uniformity, the present invention enforces the learned feature representations to meet the above requirements while fine - tuning the multi - modality model.
[0071] A) Alignment constraint:
[0072] Specifically, since there are multiple images for each class in the dataset, various image embeddings are obtained from the image encoder. To better align the text and image embeddings of a given class, the image embeddings of samples of the same class should be close to each other. To reduce the intra - class distance between the image embeddings z i and z j of the present invention, based on the L2 - norm, it is constrained that the image embeddings of samples of the same class should be close, that is, an alignment loss function is constructed
[0073]
[0074] where y i represents the true label of the image sample x i ; c represents the class serial number; represents a binary variable, that is, when y i = c, otherwise z i and z j respectively represent the image embeddings corresponding to the image samples x i and x j output by the image encoder, and || || 2 represents taking the L2 - norm.
[0075] B) Uniformity constraint:
[0076] A small distance between text (label) embeddings may lead to misclassification and make it difficult to align visual features. To address this issue, the present invention introduces a Gaussian kernel function to optimize text representation and constructs a uniformity loss function
[0077]
[0078] where exp() represents the exponential function, σ represents the width parameter of the Gaussian kernel function, t i and t j respectively represent the text m output by the text encoder i and m j corresponding text embeddings.
[0079] Based on the above analysis, the total training loss of the multimodal model of the present invention is obtained by combining the training loss of the multimodal model, the alignment loss, and the uniformity loss as follows:
[0080]
[0081] where λ represents the loss function weight.
[0082] Since soft prompts tend to prioritize knowledge of specific tasks, the model is prone to overfitting on the target task, resulting in a decline in task generalization ability. Although the present invention alleviates the overfitting problem in multimodal fine-tuning by introducing new alignment and uniformity regularization constraints, the regularization constraints are not effective for all tasks. To better adaptively adjust the regularization strength, based on the idea of meta-learning, the present invention uses the training loss gradient of each sample to adaptively adjust the strength of the regularization constraints. Specifically, by designing a weight adaptive adjustment model λ = h(z, t; θ) to predict a suitable hyperparameter λ for each sample, θ represents the parameters of the weight adaptive adjustment model. That is, the present invention uses the image embedding z and the text embedding t as the input of the weight adaptive adjustment model, so that the weight adaptive adjustment model can automatically predict the weight hyperparameter setting of the training sample.
[0083] Therefore, based on the above analysis, the present invention can train the model by constructing the following double-layer optimization problem:
[0084]
[0085] where and They represent the total loss calculated based on the training dataset and the metadata dataset, respectively. X and M represent the image set and text set in the training dataset S, respectively. Z and T represent the image embedding set and text embedding set generated based on the image set X and the text set M, respectively. Metadata sets The training dataset is used to learn the prompt parameters, while the meta dataset is used to learn the adaptive adjustment model parameters.
[0086] Based on the above principles, the present invention proposes a pre-trained multimodal model fine-tuning method based on meta-learning. Figure 2 : is a flowchart of a specific implementation of the method for fine-tuning a pre-trained multimodal model based on meta-learning of the present invention. Figure 2 As shown, the pre-trained multimodal model fine-tuning method based on meta-learning of the present invention includes the following steps:
[0087] S201: Build and pre-train a multimodal model:
[0088] Construct and pre-train a multimodal model according to actual needs, including a block embedding encoding module, an image encoder, a word embedding encoding module, and a text encoder, where:
[0089] The word embedding encoding module is used to perform word embedding encoding on the input text m to obtain word embedding and sent to the text encoder.
[0090] Text encoder is used to embed words Perform encoding processing to obtain text embedding t.
[0091] The block embedding coding module is used to perform block embedding coding on the input image x to obtain the block embedding And sent to the image encoder.
[0092] Image encoder is used to embed the block Perform encoding processing to obtain image embedding z.
[0093] S202: Constructing a data set:
[0094] According to the actual application scenarios of the multimodal model, M image samples and corresponding texts are collected to form a training dataset S, and then a meta-dataset is constructed based on the training dataset S.
[0095] In order to increase the generalization and diversity of the model, the metadata set in this embodiment The mixup data enhancement method is used to generate based on the training data set S, that is:
[0096]
[0097] Among them, represents the image sample and the corresponding text generated by the mixup data augmentation mixing method, x i , x j respectively represent two different image samples of the training dataset S, m i , m j respectively represent the texts corresponding to the image samples x i , x j ; i, j = 1, 2, …, B, where B represents the number of samples in the training dataset S, and μ represents the ratio of data augmentation mixing, and the specific value can be obtained by random sampling from the Beta distribution.
[0098] S203: Construct a meta-augmented contrastive learning model:
[0099] To fine-tune the pre-trained multi-modal model, the present invention uses the pre-trained multi-modal model as the benchmark model, replaces the word embedding encoding module with a text-side prompt module, replaces the patch embedding encoding module with an image-side prompt module, and adds a weight adaptive adjustment module, thereby constructing a meta-augmented contrastive learning model. Figure 3 is the structural diagram of the meta-augmented contrastive learning model in the present invention. As Figure 3 shown, in the meta-augmented contrastive learning model of the present invention:
[0100] The text-side prompt module is used to perform word embedding encoding on the input text m to obtain word embeddings where the parameters of the word embedding encoding adopt the parameters of the pre-trained word embedding encoding module, and the learnable text prompt vector v = [v 1 , v 2 , …, v L and the true class label y are concatenated to obtain the text prompt p = [v 1 , v 2 , …, v L , y], and then the text prompt p and the word embeddings are concatenated to obtain the text prompt embedding m′ and output it to the text encoder.
[0101] The image-side prompt module is used to perform image embedding encoding on the input image x to obtain patch embeddings where the parameters of the image embedding encoding adopt the parameters of the pre-trained image embedding encoding module, and then the learnable image prompt vector u = [u 1 , u 2 , …, u L is inserted into the image patch class token CLS and the image representation patch E obtained by the image embedding encoding to obtain the image prompt q = [CLS, u 1 , u 2 …, u L, [E], and then concatenate the image prompt q with the block embedding to obtain the image prompt embedding x' and output it to the image encoder.
[0102] The weight adaptive adjustment module is used to receive the text embedding t output by the text encoder and the image embedding z output by the image encoder, and estimate the loss function weight λ. In this embodiment, the weight adaptive adjustment model includes a multi-layer perceptron and a Sigmoid function layer, where:
[0103] The multi-layer perceptron is used to process the image embedding z and the text embedding t, and send the obtained features to the Sigmoid function layer.
[0104] The Sigmoid function layer is used to process the received features using the Sigmoid function to obtain the loss function weight λ.
[0105] According to the above structure, the forward process of the weight adaptive model in this embodiment can be expressed as:
[0106] h(z,t;θ) = σ(MLP(z,t))
[0107] where θ represents the parameters of the adaptive adjustment model, σ() represents the Sigmod function, whose role is to output the input information between 0 and 1; MLP() represents the operation of the multi-layer perceptron (MLP, Multi-Layer Perceptron).
[0108] S204: Initialize parameters:
[0109] Randomly initialize the prompt parameter w (0) = {v (0) , u (0)} and the parameters θ of the weight adaptive adjustment model (0) .
[0110] S205: Let the training round t = 1.
[0111] S206: Learn from the training dataset:
[0112] Input the training samples in the training dataset S into the meta-augmented contrastive learning model, and calculate the loss function using the following formula
[0113]
[0114] where λ i represents the loss function weight obtained by the weight adaptive adjustment module for the i-th training sample in the training dataset S; Denote the multimodal model loss function of the $i$-th training sample in the training dataset $S$, where $i = 1, 2, \ldots, B$, and $B$ represents the number of samples in the training dataset $S$. The calculation formula is as follows;
[0115]
[0116] where, denotes the probability that the image sample $x$ in the $i$-th training sample i is predicted to belong to class $c$, denotes the predicted class label of the image $x$ i ;
[0117] Denote the alignment loss function of the $i$-th training sample in the training dataset $S$. The calculation formula is as follows:
[0118]
[0119] where, $y$ i denotes the true label of the image sample $x$ in the $i$-th training sample i and $c$ represents the class serial number; denotes a binary variable, that is, when $y$ i $= c$, otherwise $z$ i and $z$ j respectively denote the image embeddings corresponding to the image samples $x$ i and $x$ j in the training dataset $S$ output by the image encoder, where $j = 1, 2, \ldots, B$; $\| \cdot \|$ 2 denotes the calculation of the L2 norm.
[0120] Denote the uniformity loss function of the $i$-th training sample in the training dataset $S$. The calculation formula is as follows:
[0121]
[0122] where, $\exp()$ represents the exponential function, $\sigma$ represents the width parameter of the Gaussian kernel function, $t$ i and $t$ j respectively denote the text embeddings corresponding to the texts $m$ i and $m$ j in the training dataset $S$ output by the text encoder.
[0123] Then calculate the gradient of the loss function and update the prompt parameter $w$ using the following formula : (t) :
[0124]
[0125] Among them, α represents a preset learning rate.
[0126] S207: Meta-dataset learning:
[0127] Input the training samples in the meta-dataset into the meta-augmented contrastive learning model, and calculate the loss function using the following formula
[0128]
[0129] where N represents the number of training samples in the meta-dataset , respectively represent the true probability and predicted probability of the training sample in the meta-dataset on class c, and n = 1, 2,..., N.
[0130] Then calculate the gradient of the loss function and update the parameters θ of the weight adaptive adjustment model using the following formula : (t) :
[0131]
[0132] where β represents a preset learning rate.
[0133] S208: Determine whether the iteration end condition is reached. If not, go to step S209; otherwise, go to step S210. The iteration end condition can be set according to actual needs. In this embodiment, the iteration end condition is to reach the preset maximum number of iterations t max .
[0134] S209: Let t = t + 1, and return to step S206.
[0135] S210: Determine the dual-prompt multi-modal model:
[0136] Determine the text prompt vector v = [v 1 , v 2 ,..., v L and the image prompt vector u = [u 1 , u 2 ,..., u L according to the prompt parameters of the current meta-augmented contrastive learning model, and remove the weight adaptive adjustment model from the meta-augmented contrastive learning model to obtain the dual-prompt multi-modal model, completing the fine-tuning of the multi-modal model.
[0137] To better illustrate the technical effects of the present invention, specific examples are used to conduct experimental verification on the present invention. In this experimental verification, two commonly used datasets, Caltech101 and FlowerS202, are fine-tuned using a common pre-trained multi-modal model. Experiments on the model generalization ability of basic class - basic class and basic class - new class are carried out on these two datasets. Among them, training prompts are set with 16 shots (16 images for each class) on the basic class, and the performance of the prompting method on the basic class and the new class is evaluated. In this setting, the model cannot see new classes during the training phase.
[0138] Three comparison methods are set in this experimental verification, namely Zero-CLIP (zero-shot CLIP), CoOp (Context Optimization CLIP), and CoCoOp (Conditional Context Optimization CLIP).
[0139] In this experimental verification, the method of the present invention is implemented using PyTorch and trained on an NVIDIA RTX3090 GPU. The backbone networks used in the present invention all adopt the CLIP network architecture, and the meta-feature enhancer uses a multi-layer MLP network.
[0140] Table 1 is a statistical table of the classification accuracy rates of the present invention and the comparison methods in different tasks in this embodiment.
[0141]
[0142] Table 1
[0143] As can be seen from the results in Table 1, the present invention has achieved the best results in the two generalization tasks of the two datasets, Caltech101 and FlowerS202, thus verifying the effectiveness of the present invention.
[0144] Although the illustrative specific embodiments of the present invention are described above for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A meta-learning-based pre-trained multimodal model feature uniform alignment method, characterized in that: The following steps are involved: S1: Build and pre-train a multimodal model according to actual needs, including a word embedding encoding module, a text encoder block, an embedding encoding module, and an image encoder, where: The word embedding encoding module is used to perform word embedding encoding on the input text m to obtain word embedding and sent to the text encoder; Text encoder is used to embed words Perform encoding processing to obtain text embedding t; The block embedding coding module is used to perform block embedding coding on the input image x to obtain the block embedding And send to the image encoder; Image encoder is used to embed the block Perform encoding processing to obtain image embedding z; S2: According to the actual application scenarios of the multimodal model, M image samples and corresponding texts are collected to form a training dataset S, and then a meta-dataset is constructed based on the training dataset S. S3: Using the pre-trained multimodal model as the baseline model, the text-side prompt module is used to replace the word embedding encoding module, the image-side prompt module is used to replace the block embedding encoding module, and a weight adaptive adjustment module is added to construct a meta-enhanced contrastive learning model, where: The text-side prompt module is used to perform word embedding encoding on the input text m to obtain word embedding The parameters of the word embedding encoding are the parameters of the pre-trained word embedding encoding module, and then the learnable text prompt vector v = [v1, v2, ..., v L ] and the true category label y to obtain the text prompt p = [v1, v2, ..., v L ,y], and then embed the text prompt p with the word The concatenated text prompt embedding m′ is output to the text encoder; The image end prompt module is used to perform image embedding encoding on the input image x to obtain block embedding The parameters of the image embedding coding are the parameters of the pre-trained image embedding coding module, and then the learnable image prompt vector u = [u1, u2, ..., u L ], and get the image prompt q = [CLS,u1,u2…,u L ,E], and then embed the image prompt q and the block The concatenated image cue embedding x′ is output to the image encoder; The weight adaptive adjustment module is used to receive the text embedding t output by the text encoder and the image embedding z output by the image encoder, and estimate the loss function weight λ; S4: Randomly initialize the prompt parameter w (0) = {v (0) ,u (0) }, weight adaptive adjustment model parameters θ (0) ; S5: Let training round t = 1; S6: Input the training samples in the training data set S into the meta-enhanced contrastive learning model and calculate the loss function using the following formula Among them, λ i represents the loss function weight obtained by the weight adaptive adjustment module for the i-th training sample in the training data set S; represents the multimodal model loss function of the i-th training sample in the training dataset S, i = 1, 2, ..., B, B represents the number of samples in the training dataset S, and the calculation formula is as follows; in, Represents the image sample x in the i-th training sample i Predict the probability of belonging to category c, Represents image x i The predicted class label of represents the alignment loss function of the i-th training sample in the training data set S, and the calculation formula is as follows: Among them, y i Represents the image sample x in the i-th training sample i The true label, c represents the category number; Represents a binary variable, that is, when y i = c, otherwise z i and z j Respectively represent the image samples x in the training dataset S output by the image encoder i and x j The corresponding image embedding, j = 1, 2, ..., B; || || 2 means to obtain the L2 norm; represents the uniformity loss function of the i-th training sample in the training data set S, and the calculation formula is as follows: Where exp( ) represents the exponential function, σ represents the width parameter of the Gaussian kernel function, t i and t j Respectively represent the text m in the training dataset S output by the text encoder i and m j The corresponding text embedding; Then calculate the loss function Gradient And use the following formula to update the prompt parameter w (t) : Among them, α represents the preset learning rate; S7: Metadata set The training sample input element in the enhanced contrast learning model uses the following formula to calculate the loss function Where N represents the metadata set The number of training samples, Metadata sets Training samples True and predicted probabilities for category c, n=1,2,…,N; Then calculate the loss function Gradient And use the following formula to update the weight adaptive adjustment model parameter θ (t) : Among them, β represents the preset learning rate; S8: Determine whether the iteration end condition is met, if yes, go to step S9, otherwise go to step S10; S9: Set training round t=t+1, and return to step S6; S10: Determine the text prompt vector v = [v1, v2, ..., v L ] and image hint vector u=[u1,u2,…,u L ], the weight adaptive adjustment model is removed from the meta-enhanced contrastive learning model to obtain the dual-cue multimodal model and complete the multimodal model fine-tuning.
2. The method according to claim 1, characterized in that The metadata set in step S2 The mixup data enhancement method is used to generate based on the training data set S.
3. The method according to claim 1, characterized in that The weight adaptive adjustment model in step S3 includes a multi-layer perceptron and a Sigmoid function layer, wherein: The multi-layer perceptron is used to process the image embedding z and the text embedding t, and the obtained features are sent to the Sigmoid function layer; The Sigmoid function layer is used to process the received features using the Sigmoid function to obtain the loss function weight λ.
Citation Information
Patent Citations
Multi-modal fine-grained sentiment analysis method based on momentum contrast learning
CN117435732A
Black box pre-training multi-modal model fine tuning method based on efficient gradient approximation
CN117668504A