Multi-mode prompt learning system and method based on data value optimization

Through the multimodal prompt learning system optimized by data value, the Shapley value is used to evaluate the contribution of image areas, filter out high-value areas and generate guided visual information, which solves the problem of inaccurate feature alignment between modals in the existing methods, and improves the classification accuracy and generalization ability of the model.

CN120429754APending Publication Date: 2025-08-05HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510596591.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing multimodal prompt learning methods fail to fully combine image semantic information in image classification tasks, resulting in inaccurate feature alignment between modals and insufficient generalization capabilities of model, especially in complex images, it is difficult to accurately capture key area information.

Method used

By building a multimodal prompt learning system, using data value optimization strategy, using Shapley values to evaluate the contribution of image areas, filter out the image area combination with the highest data value, generate guided visual information, and fuse it with the initial text prompt, optimize multimodal feature alignment, and improve the connection between the image and text modality.

Benefits of technology

It improves the efficiency of visual information utilization, enhances the classification accuracy and generalization ability of the model, strengthens the connection between image and text modes, and improves the performance of the model in image classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429754A_ABST
    Figure CN120429754A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal prompt learning system and method based on data value optimization, and belongs to the technical field of multi-modal learning. In order to improve the utilization efficiency of visual information, the method comprises the steps of extracting features of an input image; constructing a multi-mode prompt generation module, and generating an initial text prompt and an initial visual prompt; constructing a data value screening module, calculating Shapley values of different regions of the input image, screening out an image region combination with the highest data value, and generating guiding visual information; constructing a multi-modal prompt fusion module to obtain an optimized complete deep text prompt; constructing a multi-modal feature alignment module, calculating a feature similarity matrix of the image features and the text features, and guiding the image features and the text features of the input image to realize feature alignment in the shared embedding space; and a classification prediction module is constructed, category probability distribution is generated according to the aligned feature similarity matrix, final category prediction is realized by comparing probabilities of different categories, and a classification result is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal learning, and specifically relates to a multimodal prompt learning system and method based on data value optimization. Background Art

[0002] In the field of multimodal learning, vision-language models (VLMs) have been widely used in tasks such as image classification, object detection, and cross-modal retrieval. Prompt learning, as a highly efficient model adaptation method, improves the model's adaptability to downstream tasks by optimizing prompt content without changing the pre-trained model structure. In image classification tasks, the similarity between image features and text prompts generated based on category labels is typically calculated, and the best matching category is selected in the shared feature space as the image classification result.

[0003] Current multimodal prompt learning methods often rely on manually constructed (e.g., "a photo of a [CLASS]") or randomly generated text prompts. While effective in some scenarios, this approach fails to fully incorporate image semantic information, resulting in non-targeted prompts and insufficient model generalization. This is particularly true for complex images, where the model struggles to accurately capture key image regions, leading to inaccurate inter-modal feature alignment. Furthermore, most existing methods perform feature matching based on global image features, ignoring the differential contributions of different regions to classification. This can cause the model to focus on redundant or irrelevant information, impacting classification performance. Summary of the Invention

[0004] The problem to be solved by the present invention is to improve the utilization efficiency of visual information, strengthen the connection between image and text modalities, and propose a multimodal prompt learning system and method based on data value optimization.

[0005] To achieve the above object, the present invention is implemented through the following technical solutions:

[0006] A multimodal prompt learning method based on data value optimization includes the following steps:

[0007] S1. Collect input images and extract features of the input images;

[0008] S2. Construct a multimodal prompt generation module to generate an initial text prompt and an initial visual prompt based on the features of the input image obtained in step S1;

[0009] S3. Build a data value screening module to calculate the Shapley values of different regions of the input image, screen out the image region combinations with the highest data value, and generate guiding visual information;

[0010] S4. Construct a multimodal prompt fusion module to interact the guiding visual information obtained in step S3 with the initial text prompt obtained in step S2 to generate a deep text prompt guided by the visual information, which is then combined with the initial contextual prompt to obtain an optimized complete deep text prompt;

[0011] S5. Construct a multimodal feature alignment module to calculate a feature similarity matrix between image features and text features based on the optimized complete deep text hint obtained in step S4, and guide the image features and text features of the input image to achieve feature alignment in a shared embedding space;

[0012] S6. Construct a classification prediction module, generate category probability distribution based on the aligned feature similarity matrix, achieve final category prediction by comparing the probabilities of different categories, and output the classification results.

[0013] Furthermore, the specific implementation method of step S1 is to use the visual encoder of the pre-trained universal vision-language model CLIP to extract the feature representation of the input image. Among them, F image Represents the feature representation of batch input images, B represents the batch size, C represents the number of channels, H and W represent the height and width of the input image respectively, represents a real matrix.

[0014] Furthermore, the specific implementation method of step S2 of constructing the multimodal prompt generation module includes the following steps:

[0015] S2.1. Set the category list of input image Cls = [cls1, cls2, ..., cls n_cls ], where n_cls represents the number of categories in the dataset involved in the training. Cls is linked to the fixed prompt template to obtain the initial input of the text branch, which is then passed into the text encoder of the pre-trained universal vision-language model CLIP to obtain the embedded feature representation after word segmentation. Among them, L represents the number of tokens after the template is segmented, d t Represents the dimension of text embedding;

[0016] S2.2. According to the preset number of context clue words n_ctx, the middle continuous n_ctx codes are intercepted from the embedded feature representation E after word segmentation obtained in step S2.1 as the initial context clues

[0017] S2.3. According to the preset prompt depth prompts_depth, randomly initialize the composite context prompt P of prompts_depth-1 layer context ={P1,P2,…,P prompts_depth-1}, where the context hint P at each layer i In n_ctx×d t The dimension follows a normal distribution with a mean of 0 and a standard deviation of 0.02. The calculation formula is as follows:

[0018]

[0019] in, represents the normal distribution;

[0020] S2.4. Combine the initial context hint P0 obtained in step S2.2 with the composite context hint P obtained in step S2.3 context Combine to get the initial text prompt P text ={P0,P1,P2,…,P prompts_depth-1};

[0021] S2.5. Design a projection layer to convert the embedding representation of the text modality into the feature representation of the visual modality. The projection transformation matrix of each layer is Among them, d t represents the embedding dimension of CLIP’s text encoder, d v Represents the feature dimension of CLIP’s visual encoder;

[0022] S2.6. The initial text prompt P obtained in step S2.4 text After the projection layer transformation, the initial visual prompt P is obtained vision , the calculation formula is as follows:

[0023] P vision =P text W proj ;

[0024] S2.7. Initial text prompt P text and the initial visual cue P vision As the initial multimodal prompt, they jointly participate in constructing the deep structure of the image encoder and text encoder. The deep image encoder includes 12 layers of image encoding subnetworks, and each layer of visual encoding subnetwork adopts the ViT model. The deep text encoder includes 12 layers of text encoding subnetworks, and each layer of text encoding subnetwork adopts the Transformer model.

[0025] Furthermore, the specific implementation method of step S3 to construct the data value screening module includes the following steps:

[0026] S3.1. Divide the input image into blocks according to the number of custom image blocks n_patch to obtain the region combination X = {x1, x2, ..., x n_patch}, where x i represents the i-th image block;

[0027] S3.2. Use Shapley value to evaluate the contribution of each image block to the overall image representation V(S∪{x i})-V(S), where S represents the combination of selected image regions, and V(S) represents the predicted probability value of the target category in the prediction result obtained after S participates in category prediction;

[0028] For the selection of target categories, during the training process, the true category label of the image is used as the target category. During the testing process, the category label is predicted based on the global features, and the predicted category is used as the target category.

[0029] S3.3. The Monte Carlo algorithm is used to reduce the computational cost of the Shapley value, and the marginal contribution is estimated approximately by randomly sampling subsets. The i-th image block x i The calculation formula of the Shapley value is as follows:

[0030]

[0031] Among them, φ i represents the Shapley value of the i-th image block, M represents the number of random sampling groups, S j represents any subset of the region combination belonging to the input image that does not include the i-th image block;

[0032] S3.4. Calculate the Shapley value of all image blocks Select the first n image blocks with the highest value to form the candidate set S n =argmax (k) φ i ,i=1,2,…n,then from S n Select k image blocks with the highest value and construct the initial filtered set S k =S n [:k];

[0033] S3.5. From Extract the initial filtered set S k Shapley value Find the current S k The weakest block x in weakest =argminφ i ,x weakest ∈S k , in the candidate set S n Find the strongest replacement block x among the remaining candidate blocks strongest =argmaxφ j ,x strongest ∈S n \S k, temporarily perform the replacement: x weakest →x strongest , calculate the new set after temporary replacement Shapley value if If the replacement is beneficial, the replacement is performed to obtain a new filtered set S k ′=(S k \{x weakest})∪{x strongest};

[0034] S3.6. Perform a preset maximum number of local optimization iterations (max_iter rounds) until the filtered set no longer changes after a certain iteration. Optimization is then stopped, resulting in the final filtered set.

[0035] S3.7. Pass the image blocks in the image region combination X into the visual encoder of the CLIP model to obtain the feature representation F corresponding to X g ={F g1 ,F g2 ,…F gn_patch},g∈{1,2,…B}, where B represents the batch size of the input images;

[0036] S3.8. Based on the index of the image block in the final filtered set obtained in step S3.6, obtain the feature representation F of the image block in the final filtered set patch ={F g1 ,F g2 ,…F gk},F patch ∈F g , extract the Shapley value of the image block of the final filtered set F patch The features of the k image blocks are weighted and aggregated to generate guiding visual information. The calculation formula is as follows:

[0037]

[0038] Among them, F vision_guide represents the guiding visual information aggregated from the features of k image blocks, F gi Represents the feature representation of the filtered i-th image block of the input g-th image, φ gi Representative feature F gi The Shapley value of the image block pointed to.

[0039] Furthermore, the specific implementation method of step S4 of constructing the multimodal prompt fusion module includes the following steps:

[0040] S4.1. Calculate the guiding visual information F obtained in step S3vision_guide and the initial context hint P obtained in step S2 context The guided weight mapping relationship is calculated as follows:

[0041] A=softmax(F vision_guide W t P context )

[0042] Among them, A represents the guiding weight, Its size is obtained by normalizing the Shapley value of the image block area, W t It is a learnable transformation matrix, and softmax represents exponential normalization of the input value so that the sum of its weights is 1;

[0043] S4.2. Using guided weights A to fuse F vision_guide and P context , generating deep textual hints P guided by visual information guide_prompt , strengthen the guiding ability of key visual areas in multimodal prompt generation, and the calculation formula is as follows:

[0044] P guide_prompt =A T F vision_guide +P context

[0045] Among them, A T is the transpose of the guiding weights;

[0046] S4.3. Combine the initial context hint P0 with P guide_prompt Combined, we get the optimized complete deep text prompt P text_v , the expression is:

[0047] P text_v ={P0,P guide_prompt}.

[0048] Furthermore, the specific implementation method of step S5 to construct the multimodal feature alignment module includes the following steps:

[0049] S5.1. Combine the word segmentation embedding feature representation obtained in step S2 with the optimized complete deep text prompt obtained in step S4 and pass them together into the deep text encoder to obtain the corresponding deep text embedding representation. The calculation formula is:

[0050] t h =text_encoder(T h )

[0051] Among them, t h Represents the deep embedding representation of the h-th input text, T hRepresents the joint representation of the embedded feature representation of the h-th input text and the complete deep text prompt, and text_encoder represents the deep text encoder of the CLIP model frozen parameters;

[0052] S5.2. Combine the feature representation of the input image obtained in step S1 with the initial visual cue from step S2 and pass them together into the deep image encoder to obtain the corresponding deep image embedding representation, calculated as:

[0053] v g =image_encoder(V g )

[0054] Among them, v g represents the deep embedding representation of the g-th input image, V g represents the joint representation of the feature representation of the g-th input image and the initial visual cue, and image_encoder represents the deep image encoder with frozen parameters of the CLIP model;

[0055] S5.3. Calculation of t h With v g The similarity matrix between Where N is the number of samples in the batch, calculated as follows:

[0056]

[0057] Among them, θ gh v g and t h The angle between them, ‖‖·‖‖ represents the L2 norm of the vector;

[0058] S5.4. Calculate the contrast loss function based on the similarity matrix obtained in step S5.3 to optimize the feature alignment between the modalities. The calculation formula is as follows:

[0059]

[0060] Among them, S gg represents the similarity matrix between the g-th input image and the g-th input text, S gh represents the similarity matrix between the g-th input image and the h-th input text, τ represents the adjustable temperature coefficient, represents the normalization term, which calculates the similarity between the image and all texts;

[0061] S5.5. Optimize the similarity between the image and text of the input image by minimizing the loss function Loss, and achieve alignment of image features and text features in the shared embedding space.

[0062] Furthermore, the specific implementation method of step S6 to construct the classification prediction module includes the following steps:

[0063] S6.1. Use softmax to convert the similarity matrix obtained in step S5 into the probability distribution of the category. The calculation formula is as follows:

[0064]

[0065] Among them, p(y=r|v g ) represents the category probability distribution of the g-th input image, n_cls represents the total number of categories, Represents the normalized features of the r-th category text, represents the transpose of the normalized g-th input image feature vector, Represents the normalized features of the h-th category text;

[0066] S6.2. Select the category with the highest probability as the final predicted category of the image, calculated as follows:

[0067]

[0068] in, Represents the final predicted category result of image g.

[0069] A multimodal prompt learning system based on data value optimization is implemented based on the multimodal prompt learning method based on data value optimization, and includes an image preprocessing module, a multimodal prompt generation module, a data value screening module, a multimodal prompt fusion module, a multimodal feature alignment module, and a classification prediction module;

[0070] The image preprocessing module, the multimodal prompt generation module, the data value screening module, the multimodal prompt fusion module, the multimodal feature alignment module and the classification prediction module are connected in sequence;

[0071] The image preprocessing module is used to extract image features;

[0072] The multimodal prompt generation module is used to generate an initial text prompt and an initial visual prompt;

[0073] The data value screening module is used to screen the image area with the highest data value according to the Shapley value;

[0074] The multimodal prompt fusion module is used to combine the screened high-value image areas with the initial text prompts to generate deep text prompts guided by visual information;

[0075] The multimodal feature alignment module is used to align image and text features, optimize feature representation in the shared embedding space, and improve model generalization capability;

[0076] The classification prediction module is used to generate category probabilities based on the similarity between image and text features and to perform final category prediction.

[0077] Beneficial effects of the present invention:

[0078] The present invention describes a multimodal prompt learning method based on data value optimization, which proposes a data value screening mechanism, rationally evaluates the image value and designs a screening strategy to obtain the region with the highest value. Specifically, the present invention divides the input image into several regions, calculates the data value of each region as an evaluation index, and reflects the contribution of each region of the image to the overall classification task. The Shapley value in cooperative game theory provides theoretical support for our evaluation, which aims to measure the marginal contribution of an individual to the overall benefit in collaborative cooperation. In order to reduce the computational overhead while ensuring the screening effect, the present invention further proposes a screening strategy of high-value pre-screening and local iterative optimization. First, high-value image regions are pre-screened as candidate combinations, and then a target number of regions are selected from them for initialization. Based on the local replacement and value enhancement criteria, the selected region combinations are dynamically optimized, and finally the image regions that contribute most to the classification task are determined.

[0079] The present invention proposes a multimodal cue learning method based on data value optimization, which proposes a multimodal cue generation mechanism driven by guiding visual information. High-value image regions selected for prediction are weighted and aggregated to generate guiding visual information, which is then integrated with the initial textual cue. Specifically, a guiding weight mapping between visual and textual information is constructed to strengthen the guiding role of high-value image regions in textual cue generation, enabling the generation of deeper multimodal cues that are more expressive and task-relevant.

[0080] The present invention describes a multimodal prompt learning method based on data value optimization. The optimized multimodal prompt further enhances inter-modal contrast learning capabilities, strengthens the collaborative expression of image and text features in a shared embedding space, and further optimizes feature representation in combination with contrast loss, effectively improving the model's classification accuracy and generalization capabilities. The present invention introduces a data value assessment mechanism into the prompt learning framework, achieving an effective fusion of the two, effectively improving the utilization efficiency of visual information, strengthening the connection between image and text modalities, and enhancing the model's interpretability and decision transparency, further improving the model's performance in image classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 This is a schematic diagram of the structure of a multimodal prompt learning system based on data value optimization according to the present invention;

[0082] Figure 2 This is a flowchart of a multimodal prompt learning method based on data value optimization according to the present invention;

[0083] Figure 3 Schematic diagram of the data value assessment calculation method according to the present invention;

[0084] Figure 4 Schematic diagram of the high-value pre-screening and local optimization strategy of the present invention;

[0085] Figure 5 This is the overall framework diagram of the multimodal prompt learning framework based on data value optimization described in the present invention. DETAILED DESCRIPTION

[0086] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present invention and are not intended to limit the present invention. That is, the specific embodiments described herein are only some embodiments of the present invention, not all embodiments. Generally, the components of the specific embodiments of the present invention described and illustrated in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0087] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely representative of selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0088] In order to further understand the content, features and effects of the present invention, the following specific embodiments are given as examples, and the attached Figure 1 -Attached Figure 5 The detailed instructions are as follows:

[0089] Example 1:

[0090] A multimodal prompt learning method based on data value optimization includes the following steps:

[0091] S1. Collect input images and extract features of the input images;

[0092] Furthermore, the specific implementation method of step S1 is to use the pre-trained universal vision-language model CLIP visual encoder to extract the feature representation of the input image. Among them, F imageRepresents the feature representation of batch input images, B represents the batch size, C represents the number of channels, H and W represent the height and width of the input image respectively, represents a real matrix;

[0093] S2. Construct a multimodal prompt generation module to generate an initial text prompt and an initial visual prompt based on the features of the input image obtained in step S1;

[0094] Furthermore, the specific implementation method of step S2 of constructing the multimodal prompt generation module includes the following steps:

[0095] S2.1. Set the category list of input image Cls = [cls1, cls2, ..., cls n_cls ], where n_cls represents the number of categories in the dataset involved in the training. Cls is linked to the fixed prompt template to obtain the initial input of the text branch, which is then passed into the text encoder of the pre-trained universal vision-language model CLIP to obtain the embedded feature representation after word segmentation. Among them, L represents the number of tokens after the template is segmented, d t Represents the dimension of text embedding;

[0096] Furthermore, C is linked to the fixed prompt template “aphoto ofa” to obtain the initial input of the text branch;

[0097] S2.2. According to the preset number of context clue words n_ctx, the middle continuous n_ctx codes are intercepted from the embedded feature representation E after word segmentation obtained in step S2.1 as the initial context clues

[0098] S2.3. According to the preset prompt depth prompts_depth, randomly initialize the composite context prompt P of prompts_depth-1 layer context ={P1,P2,…,P prompts_depth-1}, where the context hint P at each layer i In n_ctx×d t The dimension follows a normal distribution with a mean of 0 and a standard deviation of 0.02. The calculation formula is as follows:

[0099]

[0100] in, represents the normal distribution;

[0101] S2.4. Combine the initial context hint P0 obtained in step S2.2 with the composite context hint P obtained in step S2.3 context Combine to get the initial text prompt P text={P0,P1,P2,…,P prompts_depth-1};

[0102] S2.5. Design a projection layer to convert the embedding representation of the text modality into the feature representation of the visual modality. The projection transformation matrix of each layer is Among them, d t represents the embedding dimension of CLIP’s text encoder, d v Represents the feature dimension of CLIP’s visual encoder;

[0103] S2.6. The initial text prompt P obtained in step S2.4 text After the projection layer transformation, the initial visual prompt P is obtained vision , the calculation formula is as follows:

[0104] P vision =P text W proj ;

[0105] S2.7. Initial text prompt P text and the initial visual cue P vision As the initial multimodal prompt, they jointly participate in constructing the deep structure of the image encoder and text encoder. The deep image encoder includes 12 layers of image encoding subnetworks, and each layer of visual encoding subnetwork adopts the ViT model. The deep text encoder includes 12 layers of text encoding subnetworks, and each layer of text encoding subnetwork adopts the Transformer model.

[0106] S3. Build a data value screening module to calculate the Shapley values of different regions of the input image, screen out the image region combinations with the highest data value, and generate guiding visual information;

[0107] Furthermore, the specific implementation method of step S3 to construct the data value screening module includes the following steps:

[0108] S3.1. Divide the input image into blocks according to the number of custom image blocks n_patch to obtain the region combination X = {x1, x2, ..., x n_patch}, where x i represents the i-th image block;

[0109] S3.2. Use Shapley value to evaluate the contribution of each image block to the overall image representation V(S∪{x i})-V(S), where S represents the combination of selected image regions, and V(S) represents the predicted probability value of the target category in the prediction result obtained after S participates in category prediction;

[0110] For the selection of target categories, during the training process, the true category label of the image is used as the target category. During the testing process, the category label is predicted based on the global features, and the predicted category is used as the target category.

[0111] S3.3. The Monte Carlo algorithm is used to reduce the computational cost of the Shapley value, and the marginal contribution is estimated approximately by randomly sampling subsets. The i-th image block x i The calculation formula of the Shapley value is as follows:

[0112]

[0113] Among them, φ i represents the Shapley value of the i-th image block, M represents the number of random sampling groups, S j represents any subset of the region combination belonging to the input image that does not include the i-th image block;

[0114] S3.4. Calculate the Shapley value of all image blocks Select the first n image blocks with the highest value to form the candidate set S n =argmax (k) φ i ,i=1,2,…n,then from S n Select k image blocks with the highest value and construct the initial filtered set S k =S n [:k];

[0115] S3.5. From Extract the initial filtered set S k Shapley value Find the current S k The weakest block x in weakest =argminφ i ,x weakest ∈S k , in the candidate set S n Find the strongest replacement block x among the remaining candidate blocks strongest =argmaxφ j ,x strongest ∈S n \S k , temporarily perform the replacement: x weakest →x strongest , calculate the new set after temporary replacement Shapley value if If the replacement is beneficial, the replacement is performed to obtain a new filtered set S k ′=(S k \{xweakest})∪{x strongest};

[0116] S3.6. Perform a preset maximum number of local optimization iterations (max_iter rounds) until the filtered set no longer changes after a certain iteration. Optimization is then stopped, resulting in the final filtered set.

[0117] S3.7. Pass the image blocks in the image region combination X into the visual encoder of the CLIP model to obtain the feature representation F corresponding to X g ={F g1 ,F g2 ,…F gn_patch},g∈{1,2,…B}, where B represents the batch size of the input images;

[0118] S3.8. Based on the index of the image block in the final filtered set obtained in step S3.6, obtain the feature representation F of the image block in the final filtered set patch ={F g1 ,F g2 ,…F gk},F patch ∈F g , extract the Shapley value of the image block of the final filtered set F patch The features of the k image blocks are weighted and aggregated to generate guiding visual information. The calculation formula is as follows:

[0119]

[0120] Among them, F vision_guide represents the guiding visual information aggregated from the features of k image blocks, F gi Represents the feature representation of the filtered i-th image block of the input g-th image, φ gi Representative feature F gi The Shapley value of the image block pointed to.

[0121] S4. Construct a multimodal prompt fusion module to interact the guiding visual information obtained in step S3 with the initial text prompt obtained in step S2 to generate a deep text prompt guided by the visual information, which is then combined with the initial contextual prompt to obtain an optimized complete deep text prompt;

[0122] Furthermore, the specific implementation method of step S4 of constructing the multimodal prompt fusion module includes the following steps:

[0123] S4.1. Calculate the guiding visual information F obtained in step S3 vision_guide and the initial context hint P obtained in step S2 contextThe guided weight mapping relationship is calculated as follows:

[0124] A=softmax(F vision_guide W t P context )

[0125] Among them, A represents the guiding weight, Its size is obtained by normalizing the Shapley value of the image block area, W t It is a learnable transformation matrix, and softmax represents exponential normalization of the input value so that the sum of its weights is 1;

[0126] S4.2. Using guided weights A to fuse F vision_guide and P context , generating deep textual hints P guided by visual information guide_prompt , strengthen the guiding ability of key visual areas in multimodal prompt generation, and the calculation formula is as follows:

[0127] P guide_prompt =A T F vision_guide +P context

[0128] Among them, A T is the transpose of the guiding weights;

[0129] S4.3. Combine the initial context hint P0 with P guide_prompt Combined, we get the optimized complete deep text prompt P text_v , the expression is:

[0130] P text_v ={P0,P guide_prompt}.

[0131] S5. Construct a multimodal feature alignment module to calculate a feature similarity matrix between image features and text features based on the optimized complete deep text hint obtained in step S4, and guide the image features and text features of the input image to achieve feature alignment in a shared embedding space;

[0132] Furthermore, the specific implementation method of step S5 to construct the multimodal feature alignment module includes the following steps:

[0133] S5.1. Combine the word segmentation embedding feature representation obtained in step S2 with the optimized complete deep text prompt obtained in step S4 and pass them together into the deep text encoder to obtain the corresponding deep text embedding representation. The calculation formula is:

[0134] t h =text_encoder(T h )

[0135] Among them, t h Represents the deep embedding representation of the h-th input text, T h Represents the joint representation of the embedded feature representation of the h-th input text and the complete deep text prompt, and text_encoder represents the deep text encoder of the CLIP model frozen parameters;

[0136] S5.2. Combine the feature representation of the input image obtained in step S1 with the initial visual cue from step S2 and pass them together into the deep image encoder to obtain the corresponding deep image embedding representation, calculated as:

[0137] v g =image_encoder(V g )

[0138] Among them, v g represents the deep embedding representation of the g-th input image, V g represents the joint representation of the feature representation of the g-th input image and the initial visual cue, and image_encoder represents the deep image encoder with frozen parameters of the CLIP model;

[0139] S5.3. Calculation of t h With v g The similarity matrix between Where N is the number of samples in the batch, calculated as follows:

[0140]

[0141] Among them, θ gh v g and t h The angle between them, ‖‖·|| represents the L2 norm of the vector;

[0142] S5.4. Calculate the contrast loss function based on the similarity matrix obtained in step S5.3 to optimize the feature alignment between the modalities. The calculation formula is as follows:

[0143]

[0144] Among them, S gg represents the similarity matrix between the g-th input image and the g-th input text, S gh represents the similarity matrix between the g-th input image and the h-th input text, τ represents the adjustable temperature coefficient, represents the normalization term, which calculates the similarity between the image and all texts;

[0145] S5.5. Optimize the similarity between the image and text of the input image by minimizing the loss function Loss, and achieve alignment of image features and text features in the shared embedding space.

[0146] S6. Construct a classification prediction module, generate category probability distribution based on the aligned feature similarity matrix, achieve final category prediction by comparing the probabilities of different categories, and output the classification results.

[0147] Furthermore, the specific implementation method of step S6 to construct the classification prediction module includes the following steps:

[0148] S6.1. Use softmax to convert the similarity matrix obtained in step S5 into the probability distribution of the category. The calculation formula is as follows:

[0149]

[0150] Among them, p(y=r|v g ) represents the category probability distribution of the g-th input image, n_cls represents the total number of categories, Represents the normalized features of the r-th category text, represents the transpose of the normalized g-th input image feature vector,

[0151] Represents the normalized features of the h-th category text;

[0152] S6.2. Select the category with the highest probability as the final predicted category of the image, calculated as follows:

[0153]

[0154] in, Represents the final predicted category result of image g.

[0155] Experimental analysis of the method proposed in the present invention:

[0156] This paper was experimentally validated on the Caltech101 dataset, a widely used general-purpose object dataset in computer vision. Its content covers multiple semantic categories, such as animals, everyday objects, and flowers, and exhibits good inter-class diversity. The Caltech101 dataset contains 101 object categories (including one background class), with each category containing 40 to 800 images, for a total of approximately 9,149 images.

[0157] In order to objectively evaluate the performance of the proposed method, the present invention is compared with existing prompt learning methods (CLIP, CoOp, CoCoOp and MaPLe). The base class (Base) accuracy, the unseen class (Novel) accuracy and their harmonic mean (HM) are used as evaluation indicators. Among them, the base class (Base) accuracy reflects the model's ability to recognize known categories and verifies the model's feature learning effect under sufficient training data. The unseen class (Novel) accuracy tests the model's adaptability to completely new categories and reflects the model's zero-sample or small-sample transfer performance. The harmonic mean (HM) can balance the difference in performance between the base class and the novel class, effectively suppress extreme imbalances, and avoid a single indicator dominating the evaluation. The calculation formula of HM is as follows:

[0158]

[0159] The present invention conducted experiments according to the steps described in the specific implementation method. The test results obtained are shown in Table 1. This method is represented by ShapleyPrompt, where the unit of the Base class test results and the Novel test results is accuracy (%):

[0160] Table 1 Test results of the method proposed in the present invention

[0161] Method Name Base class test results Novel test results HM test results CLIP 96.84 94.00 95.40 CoOp 98.00 89.81 93.73 CoCoOp 97.96 93.81 95.84 MaPLe 97.74 94.36 96.02 Shapley Prompt 98.82 94.79 96.76

[0162] According to the analysis of experimental results, the method proposed in this invention achieves the best performance and greatly improves the accuracy of image classification prediction and the generalization ability of the model.

[0163] Working principle of the present invention:

[0164] Feature extraction is performed on the input image to obtain basic visual representation. Initial text prompts and initial visual prompts are constructed to provide prior information suitable for the visual-language model. A data value screening mechanism is introduced to divide the input image into several regions and calculate the Shapley value of each region to measure the contribution of each image region to the overall representation. The screening strategy of high-value pre-screening and local iterative optimization is used to select the image region combination with the highest value for the classification task. The initial text prompts are optimized using the screened regions to generate deep text prompts guided by visual information. The prompts are further used for multimodal contrastive learning to guide the alignment of image and text features in a shared embedding space. The category probability distribution is generated by calculating the similarity between image and text features, and the category with the highest probability is taken to achieve image classification. The present invention improves the effectiveness of multimodal prompts through data value optimization, introduces high-value image regions to participate in the generation and optimization process of multimodal prompts, and improves classification accuracy and model generalization ability.

[0165] Example 2:

[0166] A multimodal prompt learning system based on data value optimization is implemented based on the multimodal prompt learning method based on data value optimization described in Example 1, comprising an image preprocessing module, a multimodal prompt generation module, a data value screening module, a multimodal prompt fusion module, a multimodal feature alignment module, and a classification prediction module;

[0167] The image preprocessing module, the multimodal prompt generation module, the data value screening module, the multimodal prompt fusion module, the multimodal feature alignment module and the classification prediction module are connected in sequence;

[0168] The image preprocessing module is used to extract image features;

[0169] The multimodal prompt generation module is used to generate an initial text prompt and an initial visual prompt;

[0170] The data value screening module is used to screen the image area with the highest data value according to the Shapley value;

[0171] The multimodal prompt fusion module is used to combine the screened high-value image areas with the initial text prompts to generate deep text prompts guided by visual information;

[0172] The multimodal feature alignment module is used to align image and text features, optimize feature representation in the shared embedding space, and improve model generalization capabilities;

[0173] The classification prediction module is used to generate category probabilities based on the similarity between image and text features and to perform final category prediction.

[0174] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0175] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and components may be substituted with equivalents without departing from the scope of the present application. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of these combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions within the scope of the claims.

Claims

1. A multimodal prompt learning method based on data value optimization, characterized in that: The steps include: S1. Collect input images and extract features of the input images; S2. Construct a multimodal prompt generation module to generate an initial text prompt and an initial visual prompt based on the features of the input image obtained in step S1; S3. Build a data value screening module to calculate the Shapley values of different regions of the input image, screen out the image region combinations with the highest data value, and generate guiding visual information; S4. Construct a multimodal prompt fusion module to interact the guiding visual information obtained in step S3 with the initial text prompt obtained in step S2 to generate a deep text prompt guided by the visual information, which is then combined with the initial contextual prompt to obtain an optimized complete deep text prompt; S5. Construct a multimodal feature alignment module to calculate a feature similarity matrix between image features and text features based on the optimized complete deep text hint obtained in step S4, and guide the image features and text features of the input image to achieve feature alignment in a shared embedding space; S6. Construct a classification prediction module, generate category probability distribution based on the aligned feature similarity matrix, achieve final category prediction by comparing the probabilities of different categories, and output the classification results.

2. A multimodal prompt learning method based on data value optimization according to claim 1, characterized in that: The specific implementation method of step S1 is to use the pre-trained universal vision-language model CLIP visual encoder to extract the feature representation of the input image. Among them, F image Represents the feature representation of batch input images, B represents the batch size, C represents the number of channels, H and W represent the height and width of the input image respectively, represents a real matrix.

3. A multimodal prompt learning method based on data value optimization according to claim 1 or 2, characterized in that: The specific implementation method of step S2 to construct a multimodal prompt generation module includes the following steps: S2.

1. Set the category list of input image Cls = [cls1, cls2, ..., cls n_cls ], where n_cls represents the number of categories in the dataset involved in the training. Cls is linked to the fixed prompt template to obtain the initial input of the text branch, which is then passed into the text encoder of the pre-trained universal vision-language model CLIP to obtain the embedded feature representation after word segmentation. Among them, L represents the number of tokens after template segmentation, d t Represents the dimension of text embedding; S2.

2. According to the preset number of context clue words n_ctx, the middle continuous n_ctx codes are intercepted from the embedded feature representation E after word segmentation obtained in step S2.1 as the initial context clues S2.

3. According to the preset prompt depth prompts_depth, randomly initialize the composite context prompt P of prompts_depth-1 layer context ={P1,P2,…,P prompts_depth-1 }, where the context hint P at each layer i In n_ctx×d t The dimension follows a normal distribution with a mean of 0 and a standard deviation of 0.

02. The calculation formula is as follows: in, represents the normal distribution; S2.

4. Combine the initial context hint P0 obtained in step S2.2 with the composite context hint P obtained in step S2.3 context Combine to get the initial text prompt P text ={P0,P1,P2,…,P prompts_depth-1 }; S2.

5. Design a projection layer to convert the embedding representation of the text modality into the feature representation of the visual modality. The projection transformation matrix of each layer is Among them, d t represents the embedding dimension of CLIP’s text encoder, d v Represents the feature dimension of CLIP’s visual encoder; S2.

6. The initial text prompt P obtained in step S2.4 text After the projection layer transformation, the initial visual prompt P is obtained vision , the calculation formula is as follows: P vision =P text W proj ; S2.

7. Initial text prompt P text and the initial visual cue P vision As the initial multimodal prompt, they jointly participate in constructing the deep structure of the image encoder and text encoder. The deep image encoder includes 12 layers of image encoding subnetworks, and each layer of visual encoding subnetwork adopts the ViT model. The deep text encoder includes 12 layers of text encoding subnetworks, and each layer of text encoding subnetwork adopts the Transformer model.

4. A multimodal prompt learning method based on data value optimization according to claim 3, characterized in that: The specific implementation method of step S3 to construct the data value screening module includes the following steps: S3.

1. Divide the input image into blocks according to the number of custom image blocks n_patch to obtain the region combination X = {x1, x2, ..., x n_patch }, where x i represents the i-th image block; S3.

2. Use Shapley value to evaluate the contribution of each image block to the overall image representation V(S∪{x i })-V(S), where S represents the combination of selected image regions, and V(S) represents the predicted probability value of the target category in the prediction result obtained after S participates in category prediction; For the selection of target categories, during the training process, the true category label of the image is used as the target category. During the testing process, the category label is predicted based on the global features, and the predicted category is used as the target category. S3.

3. The Monte Carlo algorithm is used to reduce the computational cost of the Shapley value, and the marginal contribution is estimated approximately by randomly sampling subsets. The i-th image block x i The calculation formula of the Shapley value is as follows: Among them, φ i represents the Shapley value of the i-th image block, M represents the number of random sampling groups, S j represents any subset of the region combination belonging to the input image that does not include the i-th image block; S3.

4. Calculate the Shapley value of all image blocks Select the first n image blocks with the highest value to form the candidate set S n =argmax (k) φ i ,i=1,2,…n,then from S n Select k image blocks with the highest value and construct the initial filtered set S k =S n [:k]; S3.

5. From Extract the initial filtered set S k Shapley value Find the current S k The weakest block x in weakest =argminφ i ,x weakest ∈S k , in the candidate set S n Find the strongest replacement block x among the remaining candidate blocks strongest =argmaxφ j ,x strongest ∈S n \S k , temporarily perform the replacement: x weakest →x strongest , calculate the new set after temporary replacement Shapley value if If the replacement is beneficial, the replacement is performed to obtain a new filtered set S' k =(S k \{x weakest })∪{x strongest }; S3.

6. Perform a preset maximum number of local optimization iterations (max_iter rounds) until the filtered set no longer changes after a certain iteration. Optimization is then stopped, resulting in the final filtered set. S3.

7. Pass the image blocks in the image region combination X into the visual encoder of the CLIP model to obtain the feature representation F corresponding to X g ={F g1 ,F g2 ,…F gn_patch },g∈{1,2,…B}, where B represents the batch size of the input images; S3.

8. Based on the index of the image block in the final filtered set obtained in step S3.6, obtain the feature representation F of the image block in the final filtered set patch ={F g1 ,F g2 ,…F gk },F patch ∈F g , extract the Shapley value of the image block of the final filtered set F patch The features of the k image blocks are weighted and aggregated to generate guiding visual information. The calculation formula is as follows: Among them, F vision_guide represents the guiding visual information aggregated from the features of k image blocks, F gi Represents the feature representation of the filtered i-th image block of the input g-th image, φ gi Representative feature F gi The Shapley value of the image block pointed to.

5. A multimodal prompt learning method based on data value optimization according to claim 4, characterized in that: The specific implementation method of step S4 to construct a multimodal prompt fusion module includes the following steps: S4.

1. Calculate the guiding visual information F obtained in step S3 vision_guide and the initial context hint P obtained in step S2 context The guided weight mapping relationship is calculated as follows: A=softmax(F vision_guide W t P context ) Among them, A represents the guiding weight, Its size is obtained by normalizing the Shapley value of the image block area, W t It is a learnable transformation matrix, and softmax represents exponential normalization of the input value so that the sum of its weights is 1; S4.

2. Using guided weights A to fuse F vision_guide and P context , generating deep textual hints P guided by visual information guide_prompt , strengthen the guiding ability of key visual areas in multimodal prompt generation, and the calculation formula is as follows: P guide_prompt =A T F vision_guide +P context Among them, A T is the transpose of the guiding weights; S4.

3. Combine the initial context hint P0 with P guide_prompt Combined, we get the optimized complete deep text prompt P text_v , the expression is: P text_v ={P0,P guide_prompt }。 6. A multimodal prompt learning method based on data value optimization according to claim 5, characterized in that: The specific implementation method of step S5 to construct a multimodal feature alignment module includes the following steps: S5.

1. Combine the word segmentation embedding feature representation obtained in step S2 with the optimized complete deep text prompt obtained in step S4 and pass them together into the deep text encoder to obtain the corresponding deep text embedding representation. The calculation formula is: t h =text_encoder(T h ) Among them, t h Represents the deep embedding representation of the h-th input text, T h Represents the joint representation of the embedded feature representation of the h-th input text and the complete deep text prompt, and text_encoder represents the deep text encoder of the CLIP model frozen parameters; S5.

2. Combine the feature representation of the input image obtained in step S1 with the initial visual cue from step S2 and pass them together into the deep image encoder to obtain the corresponding deep image embedding representation, calculated as: v g =image_encoder(V g ) Among them, v g represents the deep embedding representation of the g-th input image, V g represents the joint representation of the feature representation of the g-th input image and the initial visual cue, and image_encoder represents the deep image encoder with frozen parameters of the CLIP model; S5.

3. Calculation of t h With v g The similarity matrix between Where N is the number of samples in the batch, calculated as follows: Among them, θ gh v g and t h The angle between them, ‖‖·‖‖ represents the L2 norm of the vector; S5.

4. Calculate the contrast loss function based on the similarity matrix obtained in step S5.3 to optimize the feature alignment between the modalities. The calculation formula is as follows: Among them, S gg represents the similarity matrix between the g-th input image and the g-th input text, S gh represents the similarity matrix between the g-th input image and the h-th input text, τ represents the adjustable temperature coefficient, represents the normalization term, which calculates the similarity between the image and all texts; S5.

5. Optimize the similarity between the image and text of the input image by minimizing the loss function Loss, and achieve alignment of image features and text features in the shared embedding space.

7. A multimodal prompt learning method based on data value optimization according to claim 6, characterized in that: The specific implementation method of step S6 to construct the classification prediction module includes the following steps: S6.

1. Use softmax to convert the similarity matrix obtained in step S5 into the probability distribution of the category. The calculation formula is as follows: Among them, p(y=r|v g ) represents the category probability distribution of the g-th input image, n_cls represents the total number of categories, Represents the normalized features of the r-th category text, represents the transpose of the normalized g-th input image feature vector, Represents the normalized features of the h-th category text; S6.

2. Select the category with the highest probability as the final predicted category of the image, calculated as follows: in, Represents the final predicted category result of image g.

8. A multimodal prompt learning system based on data value optimization, implemented by the multimodal prompt learning method based on data value optimization according to claims 1-7, characterized in that: It includes image preprocessing module, multimodal prompt generation module, data value screening module, multimodal prompt fusion module, multimodal feature alignment module and classification prediction module; The image preprocessing module, the multimodal prompt generation module, the data value screening module, the multimodal prompt fusion module, the multimodal feature alignment module and the classification prediction module are connected in sequence; The image preprocessing module is used to extract image features; The multimodal prompt generation module is used to generate an initial text prompt and an initial visual prompt; The data value screening module is used to screen the image area with the highest data value according to the Shapley value; The multimodal prompt fusion module is used to combine the screened high-value image areas with the initial text prompts to generate deep text prompts guided by visual information; The multimodal feature alignment module is used to align image and text features, optimize feature representation in the shared embedding space, and improve model generalization capabilities; The classification prediction module is used to generate category probabilities based on the similarity between image and text features and to perform final category prediction.

Citation Information

Cited By

  • Image retrieval method, related device, equipment and storage medium

    CN121434431A