A multimodal image language model combining a prompt learning method and device
Patent Information
- Application Number
- CN202410477595.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-04-19
AI Technical Summary
[0006]但是,现有方法要么是将提示仅添加在文本分支上,要么将提示仅添加在图像分支上,又或者是将提示同时添加在图像和文本两个分支,仅仅通过线性映射函数进行简单的融合,并未进行更深入的交互
[0021]本公开提供的一种多模态图像语言模型结合提示学习方法、装置、设备及介质,优点在于,基于开源CLIP模型构建CCPL,冻结基础模型参数的基础上设置提示符更新模块和特征融合模块,其中,提示符更新模块的设置更新图像和文本分支的提示,鼓励图像提示符和文本提示符捕获彼此的关键信息,加强了文本和图像两个分支的提示信息之间的互动,并提高了两者捕捉对方的关键信息并将其整合到自己的能力,且在模型最后一层,设置的特征融合模块增强了输出端图像特征和文本特征的深度融合,且有效地维护了图像和文本特征的一致性,解决现有技术中提示学习方法要么只关注在单一模态内的提示信息或模态之间缺乏信息交互,泛化能力不太明显的问题,在跨模态提示学习技术较为创新并具备较高潜力。
Smart Images

Figure CN118427608B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to a multimodal image language model combined with prompting learning method, apparatus, device and medium. Background Technology
[0002] Due to the large number of parameters and limited datasets, large multimodal image language models struggle to transfer learned knowledge to downstream tasks. Combining cue learning with multimodal image language models has recently become a new research direction.
[0003] In recent years, research combining multimodal image language models and cue learning has been advancing rapidly. The overall architecture of existing mainstream cue learning methods is referenced below. Figure 1 As shown, from the perspective of model structure, there are two methods: one is to add the prompt information only to the image encoder branch or the text encoder branch (as in existing technologies). Figure 1 (a) shows the structure, and the addition of prompt information branches in the text encoder and image encoder branches (such as...). Figure 1 (b) shows the structure.
[0004] The existing CLIP (Language-Image Pre-Training) model's text input prompt is manually set to "a photo of a..." <class>The cue is assigned a class (representing the image category), and the resulting text encoder is used as the classification weight to match image features. A contrastive learning loss objective is then employed to narrow the distance between the image and its associated text description in the mapping space, eliminating mismatched images in the feature space, thereby improving model performance. However, even subtle changes to manually adjusted text cues can cause significant performance variations, requiring considerable time and effort to develop a suitable cue. Inspired by cue learning research in Natural Language Processing (NLP), the CoOp (Context Optimization) method transforms text cues into learnable context vectors. Using only a small number of labeled images, this approach achieves substantial improvements over densely adjusted manual cues on a wide range of image recognition datasets, enabling automated text cue generation for downstream tasks.
[0005] CoOp also has some shortcomings. The learned context overfits the base class, resulting in weak generalization ability on unseen classes. That is, it works well on base classes but performs poorly on unseen classes, leading to a significant drop in accuracy. To address this issue, the CoCoOp (Conditional Context Optimization) method was developed, where learnable text prompts depend on the input image. CoCoOp works by using a lightweight neural network to generate an input conditional token for each image and combining it with the learnable context vector from CoOp. This ensures that each text prompt contains information about its associated image, rather than just information about the specific category set during training. However, this method only focuses on the text branch prompt, ignoring the image branch. Visual PromptTuning (VPT) demonstrated the feasibility of this approach to address this problem. However, the intra-class variance of image features and the inter-class variance of text embeddings cause differences in data distribution, preventing a consistent performance improvement. UPT (Unified Prompt Tuning) and MaPLe (Multi-modal Prompt Learning) combine text and image cues, highlighting the advantages of multimodal cues. MaPLe generates corresponding image cues based on text through a mapping function. However, text and image cues are combined using contrastive learning, achieving interaction only through a simple linear mapping. Considering this, DCP (Deeply Coupled Cross-modal Prompt Learning) couples the text and image branches through a cross-modal cue attention module, enabling deeper interaction and thus improving performance.
[0006] However, existing methods either add hints only to the text branch, or only to the image branch, or add hints to both the image and text branches simultaneously, simply merging them through a linear mapping function without any deeper interaction. Summary of the Invention
[0007] To overcome the problems existing in the related technologies, this disclosure provides a multimodal image language model combined with prompting learning method, apparatus, device and medium to solve the above-mentioned technical problems.
[0008] This specification provides one or more embodiments of a multimodal image language model combined with a prompt learning method to construct a dataset, including text data of concatenated text prompts and image data of concatenated image prompts;
[0009] The CCPL multimodal image language model is constructed based on the open-source CLIP model. The CCPL multimodal image language model includes a prompt update module and a feature fusion module.
[0010] The prompt update module is located between two adjacent image encoders and text encoders, and is connected to the input and output image and text information of the two adjacent image encoders and text encoders. The prompt update module uses a cross-attention mechanism to fuse the image prompts and text prompts output by the upper image encoder and text encoder to obtain fused image prompts and fused text prompts respectively. The fused image prompts and fused text prompts are then summed with the image prompts and text prompts obtained by the current image encoder and text encoder according to preset weights to obtain new image prompts and text prompts, which are then used as the input image prompts and text prompts of the lower image encoder and text encoder respectively.
[0011] The feature fusion module is used to deeply fuse the image features output by the last layer image encoder and the text features output by the text encoder, and predict the classification probability.
[0012] The dataset is input into the multimodal image language model for training until the convergence condition is met, thus obtaining the trained multimodal image language model.
[0013] This specification provides one or more embodiments of a multimodal image-language model combined with prompting learning device, including:
[0014] The dataset building module is used to build datasets, including text data for concatenating text prompts and image data for concatenating image prompts;
[0015] The model building module is used to build the CCPL multimodal image language model based on the open-source CLIP model. The CCPL multimodal image language model includes a prompt update module and a feature fusion module.
[0016] The prompt update module connects to the input and output image and text information of the adjacent two layers of image encoders and text encoders. The prompt update module uses a cross-attention mechanism to fuse the image prompts and text prompts input from the upper layer image encoder and text encoder to obtain fused image prompts and fused text prompts respectively. The fused image prompts and fused text prompts are then summed with the image prompts and text prompts obtained from the current layer image encoder and text encoder according to preset weights to obtain new image prompts and text prompts, which are then used as the input image prompts and text prompts for the lower layer image encoder and text encoder respectively.
[0017] The feature fusion module is used to deeply fuse the image features output by the last layer image encoder and the text features output by the text encoder, and predict the classification probability.
[0018] The model training module is used to input the dataset into the multimodal image language model for training until the convergence condition is met, thus obtaining the trained multimodal image language model.
[0019] This specification provides a computer device in one or more embodiments, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal image-language model combined with prompting learning method as described above.
[0020] This specification provides one or more embodiments of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multimodal image-language model combined with prompting learning method described above.
[0021] This disclosure provides a multimodal image-language model combined with a cue learning method, apparatus, device, and medium. Its advantages lie in that it constructs a CCPL based on the open-source CLIP model, and sets up a cue update module and a feature fusion module on the basis of freezing the basic model parameters. The cue update module updates the cue information of the image and text branches, encouraging image and text cue messages to capture each other's key information, strengthening the interaction between the cue information of the text and image branches, and improving their ability to capture each other's key information and integrate it into their own. Furthermore, the feature fusion module in the last layer of the model enhances the deep fusion of image and text features at the output end and effectively maintains the consistency of image and text features. This addresses the problems of existing cue learning methods that either only focus on cue information within a single modality or lack information interaction between modalities, resulting in poor generalization ability. It is innovative and has high potential in cross-modal cue learning technology. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This specification provides a framework diagram of existing mainstream prompting learning methods for one or more embodiments, wherein... Figure 1 (a) A diagram showing the structure of adding prompts only to the image encoder branch or the text encoder branch. Figure 1 (b) Branch structure diagram for adding prompt information in the text encoder and image encoder branches;
[0024] Figure 2 A flowchart illustrating a multimodal image-language model combined with prompting learning method provided in one or more embodiments of this specification;
[0025] Figure 3 A schematic diagram of the architecture of a CCPL model provided for one or more embodiments of this specification;
[0026] Figure 4 Network structure diagram of the Multi-Head Attention (MHA) layer provided for one or more embodiments of this specification;
[0027] Figure 5 Experimental data tables on seven datasets for the CCPL model provided in one or more embodiments of this specification and other comparative models;
[0028] Figure 6 A comparison of the accuracy-epochs of the CCPL model and the DCP model provided for one or more embodiments of this specification on seven datasets;
[0029] Figure 7 A block diagram of a multimodal image-language model combined with prompting learning device provided for one or more embodiments of this specification;
[0030] Figure 8 This is a schematic diagram of the structure of a computer device provided for one or more embodiments of this specification. Detailed Implementation
[0031] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this invention.
[0032] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings.
[0033] Method Implementation Examples
[0034] According to embodiments of the present invention, a multimodal image-language model combined with prompting learning method is provided, such as... Figure 2 The diagram shown is a flowchart of the multimodal image-language model combined with prompting learning method provided in this embodiment. The multimodal image-language model combined with prompting learning method according to this embodiment of the invention includes:
[0035] Step S1: Obtain the dataset and construct a prompt template to generate text data including concatenated text prompts and image data including concatenated image prompts.
[0036] Step S2: Construct a CCPL (Cross-Coupled Prompt Learning) multimodal image language model based on the open-source CLIP (Contrastive Language-Image Pre-training) model. The multimodal image language model includes a prompt update module and a feature fusion module.
[0037] The prompt update module is located between the corresponding layers of the image encoder and the text encoder, and is connected to the input and output image and text information of the adjacent two layers of image encoders and text encoders. The prompt update module uses a cross-attention mechanism to fuse the image prompts and text prompts output by the upper layer image encoder and text encoder to obtain fused image prompts and fused text prompts respectively. The fused image prompts and fused text prompts are then summed with the image prompts and text prompts obtained by the current layer image encoder and text encoder according to preset weights to obtain new image prompts and text prompts, which are then used as the input image prompts and text prompts of the lower layer image encoder and text encoder respectively.
[0038] The feature fusion module is used to deeply fuse the image features output by the last layer image encoder and the text features output by the text encoder, and predict the classification probability.
[0039] In this embodiment, an optimized, open-source CLIP model is used, and the parameters of the image encoder and text encoder of the entire model are frozen, so that only the parameters of the prompt update module and the feature fusion module are updated throughout the process.
[0040] Step S3: Input the dataset into the multimodal image language model for training until the convergence condition is met, and obtain the trained multimodal image language model.
[0041] This embodiment of the method constructs a CCPL based on the open-source CLIP model. It sets up a prompt update module and a feature fusion module after freezing the basic model parameters. Specifically, image and text prompts are concatenated with the original image and text inputs as model input. The prompt update module updates the prompts in the image and text branches, encouraging them to capture each other's key information. This strengthens the interaction between the text and image prompts and improves their ability to capture and integrate each other's key information. Furthermore, the feature fusion module in the final layer of the model enhances the deep fusion of image and text features at the output end and effectively maintains the consistency of image and text features. This addresses the problem in existing prompt learning methods that either focus only on prompt information within a single modality or lack information interaction between modalities, resulting in poor generalization ability. This approach is innovative and has high potential in cross-modal prompt learning technology.
[0042] In this embodiment, the specific steps of constructing the dataset in step S1 include:
[0043] Step S11: Obtain the dataset from the public platform and divide the dataset into a training set, a validation set, and a test set. The training set and validation set are used to train and validate our CCPL model, respectively, while the test set is used to test the overall performance of the model.
[0044] In one specific embodiment, the dataset can cover a wide variety of image recognition tasks, such as classification of cars, flowers, airplanes, scenes, textures, satellite images, action videos, etc., and the dataset is divided into training set, validation set and test set according to the JSON file related to the image dataset and split.
[0045] Step S12: Preprocess the images in each dataset, such as image enhancement. Enhancement methods like random flipping, random cropping, and random rotation can be used to preprocess the images. After preprocessing, a 224×224 pixel RGB three-channel image is generated, which helps the model learn robustness to different object scales and positions. Furthermore, randomly flipping the image horizontally helps the model learn to be unaffected by the horizontal position of the object in the image, thereby improving the model's generalization ability.
[0046] Step S13: Normalize the data in each dataset; specifically,
[0047] First, calculate the mean and standard deviation of the training set. Standardizing all images in the dataset helps the model converge faster and may improve its performance. The formula is as follows:
[0048]
[0049] Where μ is the mean of the corresponding dataset, and σ is the standard deviation of the corresponding dataset.
[0050] In this embodiment, the length of both the text prompt and the image prompt is set to 16. The initial input text prompt uses the pre-trained CLIP model's "a photo of a" prompt. <class>The word embeddings are initialized, the image cues are randomly initialized from a normal distribution, and the image cues and text cues are concatenated with the original visual input and text input as the input to the model.
[0051] In this embodiment, reference Figure 3 As shown, the CCPL model framework is built based on the PyTorch deep learning library. The CCPL model backbone uses the pre-trained ViT-B / 16CLIP (Contrastive Language-Image Pre-training), which consists of 12 layers in total, including the image encoder and text encoder. The parameters of each encoder are frozen. Figure 3 A schematic diagram of a CCPL model framework with two prompt update modules is given.
[0052] In this embodiment, the prompt update module is a CCPG module consisting of two multihead attention layers, residual connections, and an MLP (Multilayer Perceptron) layer.
[0053] ①The Multihead Attention layer specifically implements the following process:
[0054] Attention score matrix calculation: One multi-head attention layer uses text cues as the query (Q) and image cues as the key (K), and performs a dot product between them to obtain a similarity matrix, which is the first attention score matrix; the other multi-head attention layer uses image cues as the query (Q) and text cues as the key (K), and performs a dot product between them to obtain a similarity matrix, which is the second attention score matrix.
[0055] Normalization: The first and second attention score matrices are normalized using the softmax activation function to obtain the probability distribution.
[0056] Weighted summation: The normalized result is mapped to the image prompt information (Vi) and the text prompt information (V), respectively. t We perform a weighted summation to obtain the final merged image prompt and merged text prompt.
[0057] Therefore, in this embodiment, at least one CCPG module is provided, and at most 11 CCPG modules are provided.
[0058] In this preferred embodiment, nine CCPG modules are preferably set up to update the text prompts and image prompts of the image encoder and text encoder of layers 1-10 of the CCPL model, respectively.
[0059] ② The residual connection is implemented in the following ways:
[0060] The image and text prompts output by the previous image encoder and text encoder are respectively output to the MLP layer with a weight coefficient of α. The fused prompts obtained by the Multihead Attention layer are updated and output to the MLP layer with a weight coefficient of (1-α). The residual connection setting can balance the importance of the prompt information of the previous encoder and the updated prompt information.
[0061] In this embodiment, the preferred value of α is 0 to 1, specifically, the value of α is 0.9.
[0062] ③ The MLP layer specifically implements the following processes:
[0063] The MLP layer adds the input image cues, text cues, fused image cues, and fused text cues using residual connections with preset weight coefficients to obtain the corresponding image cues and text cues for the next layer image encoder and text encoder.
[0064] Therefore, refer to 3 and Figure 4 As shown, where, Figure 4 The network structure diagram of the Multi-Attention (MHA) layer is shown below. The specific information processing and calculation process of the CCPG module is as follows:
[0065]
[0066] in, and These represent the text prompt and image prompt of the encoder at layer z, respectively. and These represent the text prompt and image prompt in the (z+1)th layer of the encoder, respectively. It's worth noting that the CCPG module adds these to layers 1 through N of both the text encoder and the image encoder, where N ≤ L (L represents the encoder layer number).
[0067] In this embodiment, the core principle of the Cross-Coupled Prompt Generator (CCPG) module is to update the image cues and text cues of each layer of the image encoder and text encoder using a cross-attention mechanism. This allows the image cues and text cues to capture each other's key information and integrate it into their own capabilities. This can effectively solve the problem that existing cue learning methods either only focus on cue information within a single modality or focus on cue information from both modalities but lack information interaction between modalities, resulting in poor generalization ability.
[0068] In this embodiment, the final setting of the model is a feature fusion module for deep feature fusion. Its main function is to enhance the deep fusion of image features and text features while maintaining the semantic consistency between the image and text. The feature fusion module consists of a cross-attention module, a multilayer perceptron (MLP) layer, a linear layer, and a sigmoid activation function, forming a CMF (Cross-Modal Fusion) module. This module uses the image features output from the last image encoder layer as Q in the cross-attention mechanism and the text features output from the last text encoder layer as K and V in the cross-attention mechanism. These are then input into the MLP layer, and finally, the classification probability is predicted by a classification head composed of a linear layer and a sigmoid activation function to obtain the classification result.
[0069] In this embodiment, the MLP layer also consists of two linear layers and a QuickGELU activation unit.
[0070] In this embodiment, during the CCPL model training phase, the CCPL model is optimized through a set loss function. The loss function of the CCPL model consists of the image-text matching loss function L. ITM Learning loss L compared to the original CL It consists of two parts, and their sum serves as the final optimization objective. The central idea is that the set loss function L... ITM In the Image-Text Matching (ITM) task, to determine whether image-text feature pairs match, the predicted classification probability p from the CMF module is used. itm With tag y itm The cross-entropy between them, and the original contrastive learning loss L are set. CL In contrastive learning (CL) tasks, to reduce the distance between predicted samples and positive samples while increasing the distance with negative samples, CLS tags are extracted from the last layer of the image encoder, and EOT tags are extracted from the last layer of the text encoder. The cosine similarity between the CLS and EOT tags is calculated as the prediction probability p for contrastive learning. cl Finally, the contrast learning loss is p. cl With image label y cl The cross-entropy between the predicted samples and positive samples is used to reduce the distance between them and the predicted samples, while increasing the distance between them and the negative samples. The specific formulas for the two loss functions are as follows:
[0071]
[0072] Among them, y cl With y itm All represent the true class of the image, p represents the probability of the image class prediction, D represents the probability distribution of the predicted class, and H represents the cross-entropy.
[0073] This embodiment also includes performing classification prediction on the trained CCPL model using a test set composed of few-shot image data, as detailed below:
[0074] Generate few-shot data from num_shot and save it to the shot_{num_shot}-seed_{seed_num}.pkl file, such as shot_16-seed1.pkl. The data in the pkl file is generated by randomly selecting num_shot samples from each class in the source dataset. The seed_num is 1, 2, and 3 respectively, which means that 3 pkl files are generated. Finally, the result of each num_shot is the average of these three random seeds.
[0075] For each dataset, 1 / 2 / 4 / 8 / 16-shots were selected for training.
[0076] In this embodiment, the result of num_shot is the average of these three random seeds, achieved through the following steps:
[0077] Step A1: To ensure the stability of the experimental results, three random seeds are used for each few-shot, and the final result is the average of the accuracy of seed1, seed2, and seed3.
[0078] Step A2: For seed1, initialize the number of iterations epoch = 0, the value range of epoch is 0 to 20, and the training batch size is 4;
[0079] In this embodiment, except for SUN397, 20 epochs are used for most datasets, and 5 epochs are used for 1 / 2-shot of SUN397. These parameters can optimize performance.
[0080] Step A3: The model is learned using the Stochastic Gradient Descent (SGD) optimizer, with a maximum training epoch of 20 and a learning rate of l. r =0.00001 After one warm-up period, use a learning rate l r =0.0035 for formal training, and use cosine annealing strategy to adjust the learning rate during training.
[0081] Step A4: Evaluate the CCPL model using PyTorch's built-in functions to obtain the loss function value of the training set and the results of the few-shot task. Iterate through the set epoch value. After 20 epochs of training, evaluate the current trained model using the validation set and save the experimental results throughout the process.
[0082] Step A5: After completing the seed1 experiment, take seed2 and seed3 and continue to repeat steps A1-A4. The final result is the average of the three results of seed1, seed2 and seed3.
[0083] In this embodiment, the accuracy rate is used as the evaluation metric for the few-shot task of the CCPL model. The closer the accuracy rate is to 1, the better the prediction result.
[0084] The effectiveness and advantages of this technology are illustrated below through specific examples.
[0085] Step 1: Obtain the dataset required by the model, preprocess the data, and package the data into training set, validation set and test set.
[0086] Download seven public image recognition datasets. These include StanfordCars, Flowers102, FGVCAircraft, SUN397, DTD, EuroSAT, and UCF101. These datasets cover a wide variety of image recognition tasks, such as classification of cars, flowers, airplanes, scenes, textures, satellite images, and motion videos.
[0087] Download the relevant dataset processing files, such as split_zhou_OxfordFlowers.json for the Flowers102 dataset.
[0088] The dataset was divided into training, validation, and test sets. The training and validation sets were used to train and validate our CCPL model, respectively, while the test set was used to evaluate the overall performance of the model. In this study, the dataset was divided into training, validation, and test sets based on the JSON file related to the split function in the image dataset.
[0089] Step 2: Construct a Cross-Coupled Prompt Learning model.
[0090] Set up a deep learning environment. Install the PyTorch-GPU virtual environment and PyTorch libraries on the server: python=3.8, torch=1.9.0+cu111, torchvision=0.10.0+cu111, torchaudio=0.9.0.
[0091] Install and initialize the Dassl.PyTorch utility library.
[0092] A Cross-Coupled Prompt Learning (CCPL) model framework is built based on the deep learning library PyTorch.
[0093] Step 2: Experimentation phase.
[0094] Extensive testing showed that it achieved the best average performance across seven datasets, demonstrating high accuracy in the few-shot image recognition task. Figure 5 As shown, the experimental data of the CCPL model provided in this embodiment and other comparative models are presented on seven datasets. It can be seen that our average performance on the 1, 2, 4, 8, and 16-shot datasets is improved by 1.06%, 1.29%, 1.93%, 1.74%, and 0.81% respectively compared to the second-best performing model. Furthermore, refer to... Figure 6 As shown, a comparison chart of the accuracy-epochs of the CCPL model and the DCP model of this invention on seven datasets is also provided.
[0095] Device Examples
[0096] According to embodiments of the present invention, a multimodal image-language model combined with prompting learning device is provided, such as... Figure 7 The diagram shown is a block diagram of the multimodal image-language model combined with prompting learning device provided in this embodiment. The multimodal image-language model combined with prompting learning device according to this embodiment of the invention includes:
[0097] The dataset construction module 10 is used to acquire the dataset and construct the prompt template to generate text data including concatenated text prompts and image data including concatenated image prompts.
[0098] The model building module 20 is used to build a CCPL multimodal image language model based on the open-source CLIP model. The multimodal image language model includes a prompt update module and a feature fusion module.
[0099] The prompt update module is located between the corresponding layers of the image encoder and the text encoder, and is connected to the input and output image and text information of the adjacent two layers of image encoders and text encoders. The prompt update module uses a cross-attention mechanism to fuse the image prompts and text prompts output by the upper layer image encoder and text encoder to obtain fused image prompts and fused text prompts respectively. The fused image prompts and fused text prompts are then summed with the image prompts and text prompts obtained by the current layer image encoder and text encoder according to preset weights to obtain new image prompts and text prompts, which are then used as the input image prompts and text prompts of the lower layer image encoder and text encoder respectively.
[0100] The feature fusion module is used to deeply fuse the image features output by the last layer image encoder and the text features output by the text encoder, and predict the classification probability.
[0101] The model training module 30 is used to input the dataset into the multimodal image language model for training until the convergence condition is met, and to obtain the trained multimodal image language model.
[0102] In this embodiment, the method uses a dataset construction module 10 to concatenate image and text prompts with the original image and text inputs as model input. A prompt update module updates the prompts for the image and text branches, encouraging them to capture each other's key information. This strengthens the interaction between the prompts in the text and image branches and improves their ability to capture and integrate key information from each other. Furthermore, in the final layer of the model, a feature fusion module enhances the deep fusion of image and text features at the output, effectively maintaining the consistency of image and text features. This addresses the problem in existing prompt learning methods that either focus only on prompts within a single modality or lack information interaction between modalities, resulting in poor generalization ability. This approach is innovative and has high potential in cross-modal prompt learning technology.
[0103] In this embodiment, a data processing module 40 is also provided, which is used to perform preprocessing and normalization operations on the images in each dataset.
[0104] Preprocessing includes image enhancement operations, such as random flipping, random cropping, and random rotation. After preprocessing, a 224×224 pixel RGB three-channel image is generated, which helps the model learn the robustness of objects at different scales and positions. Randomly flipping the image horizontally helps the model learn to be unaffected by the horizontal position of objects in the image, thereby improving the model's generalization ability.
[0105] The normalization process is as follows:
[0106] First, calculate the mean and standard deviation of the training set. Standardizing all images in the dataset helps the model converge faster and may improve its performance. The formula is as follows:
[0107]
[0108] Where μ is the mean of the corresponding dataset, and σ is the standard deviation of the corresponding dataset.
[0109] In this embodiment, text prompts and image prompts set by the data processing module 40 are concatenated with the visual input and text input as input to the model. The length of both text prompts and image prompts is set to 16. The initial input text prompt uses the "aphoto of a" prompt from the pre-trained CLIP model. <class>The word embedding is used for initialization, and the image prompt is randomly initialized from a normal distribution.
[0110] In this embodiment, the prompt update module is a CCPG module consisting of two multihead attention layers, residual connections, and an MLP (Multilayer Perceptron) layer.
[0111] ①The Multihead Attention layer specifically implements the following process:
[0112] Attention score matrix calculation: One multi-head attention layer uses text cues as the query (Q) and image cues as the key (K), and performs a dot product between them to obtain a similarity matrix, which is the first attention score matrix; the other multi-head attention layer uses image cues as the query (Q) and text cues as the key (K), and performs a dot product between them to obtain a similarity matrix, which is the second attention score matrix.
[0113] Normalization: The first and second attention score matrices are normalized using the softmax activation function to obtain the probability distribution.
[0114] Weighted summation: The normalized result is mapped to the image prompt information (Vi) and the text prompt information (V), respectively. t We perform a weighted summation to obtain the final merged image prompt and merged text prompt.
[0115] Therefore, in this embodiment, at least one CCPG module is provided, and at most 11 CCPG modules are provided.
[0116] In this preferred embodiment, nine CCPG modules are preferably set up to update the text prompts and image prompts of the image encoder and text encoder of layers 1-10 of the CCPL model, respectively.
[0117] ② The residual connection is implemented in the following ways:
[0118] The image and text prompts output by the previous image encoder and text encoder are respectively output to the MLP layer with a weight coefficient of α. The fused prompts obtained by the Multihead Attention layer are updated and output to the MLP layer with a weight coefficient of (1-α). The residual connection setting can balance the importance of the prompt information of the previous encoder and the updated prompt information.
[0119] In this embodiment, the preferred value of α is 0 to 1, specifically, the value of α is 0.9.
[0120] ③ The MLP layer specifically implements the following processes:
[0121] The MLP layer is used to add the image prompts, text prompts, fused image prompts, and fused text prompts from the residual connection inputs using preset weight coefficients of the residual connection, and obtain the image prompts and text prompts for the next layer image encoder and text encoder respectively.
[0122] Therefore, refer to 3 and Figure 4 As shown, where, Figure 4 The network structure diagram of the Multi-Attention (MHA) layer is shown below. The specific information processing and calculation process of the CCPG module is as follows:
[0123]
[0124] in, and These represent the text prompt and image prompt of the encoder at layer z, respectively. and These represent the text prompt and image prompt of the encoder at layer z+1, respectively.
[0125] In this embodiment, the feature fusion module is a cross-attention module, a multilayer perceptron (MLP), and a CMF (Cross-Modal Fusion) module consisting of a linear layer and a sigmoid function. This module uses the image features output by the last layer image encoder as Q in the cross-attention mechanism and the text features output by the last layer text encoder as K and V in the cross-attention mechanism. These are then input into the MLP layer and finally passed through a classification head consisting of a linear layer and a sigmoid activation function to predict the classification probability, thereby obtaining the classification result.
[0126] In this embodiment, the loss function of the CCPL model is derived from the image-text matching loss function L. ITM Learning loss L compared to the original CL It consists of two parts, and their sum serves as the final optimization objective. The core idea is that in the Image-Text Matching (ITM) task, to determine whether image-text feature pairs match, the predicted classification probability p from the CMF module is used. itm With tag y itm In contrastive learning (CL) tasks, to reduce the distance between predicted samples and positive samples while increasing the distance with negative samples, CLS tokens are extracted from the last layer of the image encoder, and EOT tokens are extracted from the last layer of the text encoder. The cosine similarity between the CLS tokens and EOT tokens is calculated as the prediction probability p for contrastive learning. cl Finally, the contrast learning loss is p. cl With image label y cl The cross-entropy between the predicted samples and positive samples is used to reduce the distance between them and the predicted samples, while increasing the distance between them and the negative samples. The specific formulas for the two loss functions are as follows:
[0127] L CL =E (I,T)~D H(y cl p cl (I,T))
[0128] L ITM =E (I,T)~D H(y itm p itm (I,T))
[0129] L = L CL +L ITM ;
[0130] Among them, y cl With y itm All represent the true class of the image, p represents the probability of the image class prediction, D represents the probability distribution of the predicted class, and H represents the cross-entropy.
[0131] In this embodiment, the model training module 30 performs classification prediction on the trained CCPL model using a test set composed of few-shot image data. Extensive experiments have demonstrated that our proposed method achieves the best overall performance on the few-shot image recognition task.
[0132] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operations of each module processing step can be understood with reference to the description of the method embodiments, and will not be repeated here.
[0133] like Figure 8 As shown, the present invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the multimodal image-language model combined with prompting learning method in the above embodiments, or when the computer program is executed by a processor, it implements the multimodal image-language model combined with prompting learning method in the above embodiments. When the computer program is executed by the processor, it implements the following method steps:
[0134] Step S1: Obtain the dataset and construct a prompt template to generate text data including concatenated text prompts and image data including concatenated image prompts.
[0135] Step S2: Construct the CCPL multimodal image language model based on the open-source CLIP model. The multimodal image language model includes a prompt update module and a feature fusion module.
[0136] Step S3: Input the dataset into the multimodal image language model for training until the convergence condition is met, and obtain the trained multimodal image language model.
[0137] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are well known to those skilled in the art.< / class> < / class> < / class>
Claims
1. A multimodal image-based language model combined with a prompting learning method, characterized in that, Including the following steps: Obtain the dataset and construct a prompt template to generate text data including concatenated text prompts and image data including concatenated image prompts; The CCPL multimodal image language model is constructed based on the open-source CLIP model. The CCPL multimodal image language model includes a prompt update module and a feature fusion module. The prompt update module connects to the input and output image and text information of the adjacent two layers of image encoders and text encoders. The prompt update module uses a cross-attention mechanism to fuse the image prompts and text prompts output by the upper layer image encoder and text encoder to obtain fused image prompts and fused text prompts respectively. The fused image prompts and fused text prompts are then summed with the image prompts and text prompts obtained by the current layer image encoder and text encoder according to preset weights to obtain new image prompts and text prompts, which are then used as the input image prompts and text prompts of the lower layer image encoder and text encoder respectively. The feature fusion module is used to deeply fuse the image features output by the last layer image encoder and the text features output by the text encoder, and predict the classification probability. The dataset is input into the multimodal image-language model for training until the convergence condition is met, thus obtaining a well-trained multimodal image-language model. The prompt update module is a CCPG module consisting of two multi-head attention layers, residual connections, and an MLP layer; The residual connection outputs the image cues and text cues from the previous image encoder and text encoder respectively to the MLP layer with a weight coefficient of α. The fused image cues and fused text cues obtained after being updated by two multi-head attention layers are output to the MLP layer with a weight coefficient of 1-α. The MLP layer adds the input image prompts, text prompts, fused image prompts, and fused text prompts using residual connections and preset weight coefficients, respectively, to obtain the image prompts and text prompts for the next layer image encoder and text encoder; the specific information processing and calculation process of the CCPG module is as follows: ; in, and These represent the text prompt and image prompt of the encoder at layer z, respectively. and These represent the text prompt and image prompt of the encoder at layer z+1, respectively.
2. The multimodal image-language model combined with prompting learning method as described in claim 1, characterized in that, The multi-head attention layer uses text prompt information as query Q and image prompt information as key K. The dot product of the two is used to obtain an attention score matrix. The attention score matrix is normalized by an activation function to obtain a probability distribution. The result of the normalization process is weighted and summed with the image prompt information V to obtain the fused image prompt. Another multi-head attention layer uses image prompt information as query Q and text prompt information as key K. The dot product of the two is used to obtain an attention score matrix. The attention score matrix is normalized by an activation function to obtain a probability distribution. The result of the normalization process is weighted and summed with the text prompt information V to obtain the fused text prompt.
3. The multimodal image-language model combined with prompting learning method as described in claim 2, characterized in that, The value of α is 0.
9.
4. The multimodal image-language model combined with prompting learning method as described in claim 1, characterized in that, The feature fusion module is a cross-attention module, an MLP layer, a linear layer, and an activation function-based CMF module. The image features output by the last layer image encoder are Q in the cross-attention mechanism, and the text features output by the last layer text encoder are K and V in the cross-attention mechanism. These are simultaneously input into the MLP layer, and then the classification probability is predicted by a classification head composed of a linear layer and a sigmoid activation function to obtain the classification result.
5. The multimodal image-language model combined with prompting learning method as described in claim 1, characterized in that, The loss function of the CCPL multimodal image-language model is composed of the image-text matching loss function L. ITM Learning loss L compared to the original CL It consists of two parts, and the specific formulas for the two loss functions are as follows: ; in, and All represent the true category of the image. p This represents the probability of predicting the image category. D H represents the probability distribution for predicted classification, and H represents the cross-entropy.
6. A multimodal image-language model combined with prompting learning device, characterized in that, include: The dataset building module is used to obtain and build the dataset, construct the prompt template, and generate text data including concatenated text prompts and image data including concatenated image prompts. The model building module is used to build the CCPL multimodal image language model based on the open-source CLIP model. The CCPL multimodal image language model includes a prompt update module and a feature fusion module. The prompt update module connects to the input and output image and text information of the adjacent two layers of image encoders and text encoders. The prompt update module uses a cross-attention mechanism to fuse the image prompts and text prompts output by the upper layer image encoder and text encoder to obtain fused image prompts and fused text prompts respectively. The fused image prompts and fused text prompts are then summed with the image prompts and text prompts obtained by the current layer image encoder and text encoder according to preset weights to obtain new image prompts and text prompts, which are used as the input image prompts and text prompts of the lower layer image encoder and text encoder respectively. The prompt update module is a CCPG module consisting of two multi-head attention layers, residual connections, and an MLP layer. The residual connection outputs the image cues and text cues from the previous image encoder and text encoder respectively to the MLP layer with a weight coefficient of α. The fused image cues and fused text cues obtained after being updated by two multi-head attention layers are output to the MLP layer with a weight coefficient of 1-α. The MLP layer adds the input image prompts, text prompts, fused image prompts, and fused text prompts using residual connections and preset weight coefficients, respectively, to obtain the image prompts and text prompts for the next layer image encoder and text encoder; the specific information processing and calculation process of the CCPG module is as follows: ; in, and These represent the text prompt and image prompt of the encoder at layer z, respectively. and These represent the text prompt and image prompt of the encoder at layer z+1, respectively; The feature fusion module is used to deeply fuse the image features output by the last layer image encoder and the text features output by the text encoder, and predict the classification probability. The model training module is used to input the dataset into the multimodal image language model for training until the convergence condition is met, thus obtaining the trained multimodal image language model.
7. The multimodal image-language model combined with prompting learning device as described in claim 6, characterized in that, The A multi-head attention layer uses text cue information as query Q and image cue information as key K. The dot product of the two is used to obtain an attention score matrix. The attention score matrix is normalized by an activation function to obtain a probability distribution. The result of the normalization process is weighted and summed with the image cue information V to obtain the fused image cue. Another multi-head attention layer uses image prompt information as query Q and text prompt information as key K. The dot product of the two is used to obtain an attention score matrix. The attention score matrix is normalized by an activation function to obtain a probability distribution. The result of the normalization process is weighted and summed with the text prompt information V to obtain the fused text prompt.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal image language model combined with prompting learning method as described in any one of claims 1 to 5.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multimodal image language model combined with prompting learning method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-modal representation learning method based on text guide image block screening
CN117421591A
False news early detection method, system, equipment and medium
CN117874607A