A large model-based image-text matching method

CN122049905BActive Publication Date: 2026-09-11北京世元科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610139286.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-09-11
Estimated Expiration
2046-02-02

AI Technical Summary

Technical Problem

然而,现有方法仍存在若干问题:首先,多数方法在对图像和文本进行融合建模时未能充分考虑两者在空间结构和语义层面的复杂依赖关系,导致对关键匹配线索的识别能力不足;其次,现有解码器结构在执行自注意力计算时对所有token均赋予相同注意范围,忽视了图像与文本之间应有的结构约束与关联路径,限制了模型的注意力机制表达能力;此外,缺乏一种能够显式引导模型关注语义一致性区域的机制,可能导致模型输出与图像实际语义之间存在偏差,影响最终匹配结果的准确性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049905B_ABST
    Figure CN122049905B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on big model's image-text matching method, including the following steps: step one: to be matched image is input to the CLIP encoder in improved BLIP-2 model;Step two: visual feature sequence is input to Q-Former module, and text label sequence is input text encoder;Step three: image semantic vector and text semantic vector are input to image-text alignment module;Step four: joint input sequence is input each layer decoder of improved LLaMA, and attention path planning matrix is generated by learning attention path planning subnetwork;Step five: introduce semantic consistency calibration mechanism;Step six: the hidden state output by the most top layer decoder is input to the output layer of improved BLIP-2 model, and image-text matching result is generated.The application effectively improves the accuracy and robustness of image-text matching, and is suitable for multi-modal intelligent retrieval, image-text question and answer and image-text relationship understanding and other complex cross-modal application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning and image-text matching technology, and in particular to an image-text matching method based on a large model. Background Technology

[0002] With the rapid development of artificial intelligence technology, image-text matching, as a key task at the intersection of computer vision and natural language processing, has gradually become one of the core technologies in application scenarios such as image question answering, image captioning generation, and cross-modal retrieval. Traditional image-text matching methods mostly rely on separate encoding methods for images and text, employing shallow feature fusion strategies to evaluate the semantic relevance between images and text. These methods have significant limitations in handling complex semantic relationships, fine-grained semantic alignment, and multimodal information fusion, making it difficult to meet the high requirements for accuracy and generalization ability in practical applications.

[0003] In recent years, with the development of deep learning technology, especially the widespread application of large-scale pre-trained models, multimodal models such as CLIP and BLIP-2 have made significant progress in image-text matching tasks. These models extract image and text features through visual encoders and text encoders respectively, and introduce multimodal alignment modules to model cross-modal semantic relationships, achieving more accurate matching results. However, existing methods still have several problems: First, most methods fail to fully consider the complex dependencies between images and text at the spatial structure and semantic levels when fusing and modeling images and text, resulting in insufficient recognition of key matching clues; second, existing decoder structures assign the same attention range to all tokens when performing self-attention calculations, ignoring the structural constraints and association paths that should exist between images and text, limiting the expressive power of the model's attention mechanism; in addition, the lack of a mechanism to explicitly guide the model to focus on semantically consistent regions may lead to a deviation between the model output and the actual semantics of the image, affecting the accuracy of the final matching result.

[0004] Therefore, how to provide a text-image matching method based on a large model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a large-model-based image-text matching method. This invention fully integrates a multimodal large-model approach, an attention path planning mechanism, and a semantic consistency calibration mechanism. It details the overall process of constructing image-text semantic vectors, performing cross-modal alignment, introducing structured attention constraints, and regressing the matching results. This method utilizes an improved BLIP-2 model to acquire semantic features of images and text. It optimizes intermodal matching quality by introducing an image-text alignment mechanism based on InfoNCE loss. A learned attention path planning sub-network is introduced into the LLaMA decoder to provide structured guidance for attention weights. Simultaneously, the semantic consistency calibration mechanism enhances the regulatory effect of image semantics on the decoder's hidden state, thereby significantly improving the accuracy and robustness of image-text matching. This invention offers advantages such as stronger semantic alignment, more accurate attention modeling, and more reliable matching results.

[0006] According to an embodiment of the present invention, a large-model-based image-text matching method includes the following steps:

[0007] Step 1: Obtain the image to be matched. Input the image to be matched into the CLIP encoder in the improved BLIP-2 model to obtain the visual feature sequence. The improved BLIP-2 model includes a CLIP encoder, a Q-Former module, a text encoder, an improved LLaMA, and an output layer, and introduces an image-text alignment module.

[0008] Step 2: Input the visual feature sequence into the Q-Former module to obtain the image semantic vector; obtain the text to be matched, divide the text to be matched into a text tag sequence, input the text tag sequence into the text encoder to obtain the text semantic vector;

[0009] Step 3: Input the image semantic vector and the text semantic vector into the image-text alignment module, calculate the alignment loss based on InfoNCE, and update the trainable parameters in the Q-Former module and the text encoder according to the alignment loss to obtain multimodal alignment features;

[0010] Step 4: Concatenate the multimodal alignment features with the embedding vectors corresponding to the text tag sequence to form a joint input sequence. Input the joint input sequence into the decoder of each layer of the improved LLaMA. Before the self-attention calculation, generate the attention path planning matrix by learning the attention path planning sub-network, and use the attention path planning matrix to perform structured constraints on the attention weights of the self-attention weight matrix of the current layer.

[0011] Step 5: After multi-head self-attention calculation, a semantic consistency calibration mechanism is introduced. The hidden state of the current layer is projected and fused according to the image semantic vector to generate gating coefficients. The hidden state of the current layer is adjusted according to the gating coefficients. The adjusted hidden state is input into the feedforward network module to generate the output of the current layer decoder.

[0012] Step 6: Input the hidden state output by the top-level decoder into the output layer of the improved BLIP-2 model, perform regression calculation on the hidden state, and generate the image-text matching result.

[0013] Optionally, step one specifically includes:

[0014] The images to be matched are scaled and normalized in size to obtain preprocessed images;

[0015] The preprocessed image is divided into N image blocks according to a preset block partitioning rule, and the image blocks are flattened to obtain image block vectors.

[0016] The image patch vector is linearly mapped to generate the image patch embedding vector, and a classification label embedding vector representing the whole image is added before the image patch embedding vector. The classification label embedding vector and the image patch embedding vector are concatenated to form the image embedding sequence.

[0017] The image embedding sequence is overlaid with position encoding. The image embedding sequence after overlay position encoding is input into the CLIP encoder in the improved BLIP-2 model. In the CLIP encoder, the image embedding sequence is extracted by a multi-head self-attention network and a feedforward network in sequence, and the corresponding visual feature sequence is output.

[0018] Optionally, step two specifically includes:

[0019] The visual feature sequence is input into the Q-Former module. M learnable query vectors are preset in the Q-Former module. The learnable query vectors are linearly transformed to obtain the query representation. The visual feature sequence is linearly transformed to obtain the key representation and value representation.

[0020] The query representation is input into a multi-head self-attention network to model the relationship between the learnable query vectors, thus obtaining the updated query representation.

[0021] The updated query representation, along with the key and value representations, is input into a multi-head cross-attention network. The correlation between the updated query representation and the visual feature sequence is calculated, and the image association representation corresponding to each learnable query vector is output.

[0022] The image association representation is subjected to a nonlinear transformation, and the image association representations are combined according to a preset method to generate an image semantic vector.

[0023] The text to be matched is divided into P text tokens according to a preset word segmentation rule, resulting in a text token sequence;

[0024] Each text token in the text token sequence is mapped to its corresponding word embedding vector, and a start token for identifying the start position of the text and an end token for identifying the end position of the text are added to the text token sequence to obtain an extended text token sequence containing the start and end tokens;

[0025] The word embedding vectors in the extended text tag sequence are superimposed with positional encoding to form a text embedding sequence with positional information;

[0026] The text embedding sequence with location information is input into the text encoder. The text embedding sequence is then processed by multi-head self-attention operation and feedforward operation to extract features and output the hidden representations corresponding to each text tag.

[0027] The hidden representations are aggregated according to a preset method to generate a text semantic vector that represents the overall semantics of the text to be matched.

[0028] Optionally, step three specifically includes:

[0029] Image semantic vectors and text semantic vectors are organized into preset batches to form image semantic vector sets and text semantic vector sets within the same batch;

[0030] In the image-text alignment module, the image semantic vector of the same image and the text semantic vector corresponding to the same image are set as image-text positive sample pairs;

[0031] Image semantic vectors and non-corresponding text semantic vectors within a batch are set as negative image-text sample pairs in the image-to-text direction;

[0032] Set the text semantic vector and the non-corresponding image semantic vector within the batch as a text-to-image negative sample pair;

[0033] The similarity between the semantic vector of each image in the batch and the semantic vector of each text in the batch is calculated to obtain the set of similarity between the image and the text.

[0034] Based on the similarity set from image to text and the preset temperature parameter, a normalization operation is performed to obtain the alignment probability distribution from image to text.

[0035] Based on the alignment probability distribution from image to text and the labeling information of positive image-text samples, calculate the alignment loss from image to text.

[0036] The similarity between each text semantic vector in the batch and each image semantic vector in the batch is calculated to obtain the set of similarity between text and image.

[0037] Based on the similarity set from text to image and the preset temperature parameter, a normalization operation is performed to obtain the alignment probability distribution from text to image.

[0038] Based on the alignment probability distribution of text to image direction and the labeling information of positive image-text sample pairs, calculate the alignment loss of text to image direction;

[0039] The alignment loss from image to text direction is summed with the alignment loss from text to image direction to obtain the InfoNCE alignment loss.

[0040] The trainable parameters in the Q-Former module and the text encoder are updated using gradients based on the InfoNCE alignment loss. The updated image semantic vector output by the Q-Former module is combined with the updated text semantic vector output by the text encoder to generate multimodal alignment features that represent the cross-modal alignment results between the image and the text.

[0041] Optionally, step four specifically includes:

[0042] The multimodal alignment features and the embedding vectors corresponding to the text tag sequence are concatenated along the feature dimension to obtain the joint input sequence;

[0043] Position encoding is superimposed on each embedding vector in the joint input sequence to obtain a joint embedding sequence with position information;

[0044] The joint embedding sequence is sequentially input into each layer of the decoder of the improved LLaMA;

[0045] In each decoder layer, before performing multi-head self-attention computation, the hidden state output by the previous decoder layer is obtained, and the hidden state is input into the learning attention path planning sub-network to generate the attention path planning matrix corresponding to the multi-head self-attention weight matrix of the current layer.

[0046] During the multi-head self-attention calculation process of the current layer, the multi-head self-attention weight matrix of the current layer is constrained element by element using the attention path planning matrix to obtain the multi-head self-attention weights constrained by the path planning.

[0047] The value representations corresponding to the key representations of the current layer decoder are weighted and summed based on the multi-head self-attention weights constrained by path planning to generate the updated hidden state of the current layer.

[0048] Optionally, the learning attention path planning sub-network includes a cross-modal feature fusion layer, a content grouping and relation calculation layer, and a path planning matrix generation layer;

[0049] The cross-modal feature fusion layer performs linear mapping between the current layer's hidden state and the image semantic vector, and performs cross-modal fusion through concatenation and gating to obtain cross-modal fused features;

[0050] The content grouping and relation calculation layer performs nonlinear transformation on the cross-modal fusion features to generate content scores. Based on the content scores, a relation graph is constructed through learnable soft clustering to characterize the relationship between the hidden states corresponding to the sequence elements in the joint input sequence.

[0051] The path planning matrix generation layer takes the relationship graph as input, linearly maps and activates the function to generate a path planning matrix with a range of element values. The path planning matrix is ​​then sparsified by sorting to obtain an attention path planning matrix used to constrain the multi-head self-attention connection relationship.

[0052] Optionally, the semantic consistency calibration mechanism specifically includes:

[0053] In the current layer decoder, the current layer hidden state after multi-head self-attention calculation and path planning constraints is obtained to obtain the current layer hidden state used for semantic consistency calibration.

[0054] Projection operations are performed on the image semantic vector to obtain a projection vector that is in the same feature space as the current layer hidden state;

[0055] The projection vector is fused with the current layer hidden state according to a preset method to form a fused representation for gating coefficient calculation;

[0056] Perform nonlinear and linear transformations on the fused representation to obtain gating coefficients corresponding to the hidden state dimension of the current layer;

[0057] The hidden state of the current layer is adjusted element by element based on the gating coefficient to obtain the calibrated hidden state after semantic consistency calibration;

[0058] The calibration hidden state is input into the feedforward network module, and linear and nonlinear transformations are performed on the calibration hidden state. The residual connection and normalization operation are then superimposed to generate the output of the current layer decoder.

[0059] Optionally, step six specifically includes:

[0060] Obtain the updated hidden state output from the top-level decoder of the improved LLaMA and use it as the hidden state input for the text-to-image matching regression.

[0061] The hidden state of the image-text matching regression is input into the output layer of the improved BLIP-2 model. In the output layer, linear and nonlinear transformations are performed on the hidden state of the image-text matching regression to obtain the intermediate matching vector used to represent the image-text matching relationship.

[0062] A regression operation is performed on the intermediate matching vector to generate an image-text matching score that characterizes the degree of matching between the image to be matched and the text to be matched.

[0063] Based on the comparison between the image-text matching score and the preset threshold, the corresponding image-text matching result is output.

[0064] The beneficial effects of this invention are:

[0065] This invention significantly improves the accuracy and robustness of image-text semantic alignment and matching by constructing a large-model-based image-text matching method, achieving several beneficial effects that are difficult to achieve with existing technologies. First, this invention introduces an improved BLIP-2 model, deeply integrating the visual encoder, Q-Former, text encoder, and decoder structures, making feature extraction of images and text more efficient and semantically expressive, fundamentally enhancing the quality of image-text semantic representation. Second, by employing an InfoNCE-based alignment loss function in the image-text alignment module, positive and negative sample image-text pairs are effectively constructed, improving the discriminative ability of cross-modal representation learning and enhancing the mapping consistency between image semantics and text semantics.

[0066] This invention introduces a learning-based attention path planning sub-network into the decoder, which structurally constrains the connections of multi-head self-attention, enabling the model to proactively plan the attention flow, strengthen the attention relationships between key regions, suppress redundant dependencies, and improve the expressive power and generalization performance of the attention mechanism. Building upon this, a semantic consistency calibration mechanism is further proposed, guiding image semantics to participate in the hidden state adjustment process of each layer of the decoder. Through gating adjustment, dynamic consistency correction of image and text semantics is achieved during deep fusion, thereby improving the semantic consistency and matching accuracy of the final output features.

[0067] This invention not only solves the problems of inaccurate image-text matching and alignment, insufficient semantic fusion, and attention diffusion in the prior art, but also provides a technical solution with a clear structure, rigorous logic, and superior effect for multimodal information processing, which has broad application prospects and promotional value. Attached Figure Description

[0068] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0069] Figure 1This is an overall flowchart of a large-model-based image-text matching method proposed in this invention;

[0070] Figure 2 This is a schematic diagram of the improved BLIP-2 model structure of the image-text matching method based on a large model proposed in this invention;

[0071] Figure 3 This is a diagram of the learning attention path planning subnetwork structure of a large-model-based image-text matching method proposed in this invention. Detailed Implementation

[0072] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0073] refer to Figure 1-3 A large-model-based image-text matching method includes the following steps:

[0074] Step 1: Obtain the image to be matched. Input the image to be matched into the CLIP encoder in the improved BLIP-2 model to obtain the visual feature sequence. The improved BLIP-2 model includes a CLIP encoder, a Q-Former module, a text encoder, an improved LLaMA, and an output layer, and introduces an image-text alignment module.

[0075] Step 2: Input the visual feature sequence into the Q-Former module to obtain the image semantic vector; obtain the text to be matched, divide the text to be matched into a text tag sequence, input the text tag sequence into the text encoder to obtain the text semantic vector;

[0076] Step 3: Input the image semantic vector and the text semantic vector into the image-text alignment module, calculate the alignment loss based on InfoNCE, and update the trainable parameters in the Q-Former module and the text encoder according to the alignment loss to obtain multimodal alignment features;

[0077] Step 4: Concatenate the multimodal alignment features with the embedding vectors corresponding to the text tag sequence to form a joint input sequence. Input the joint input sequence into the decoder of each layer of the improved LLaMA. Before the self-attention calculation, generate the attention path planning matrix by learning the attention path planning sub-network, and use the attention path planning matrix to perform structured constraints on the attention weights of the multi-head self-attention of the current layer.

[0078] Step 5: After multi-head self-attention calculation, a semantic consistency calibration mechanism is introduced. The hidden state of the current layer is projected and fused according to the image semantic vector to generate gating coefficients. The hidden state of the current layer is adjusted according to the gating coefficients. The adjusted hidden state is input into the feedforward network module to generate the output of the current layer decoder.

[0079] Step 6: Input the hidden state output by the top-level decoder into the output layer of the BLIP-2 model, perform regression calculation on the hidden state, and generate the image-text matching result.

[0080] In this embodiment, step one specifically includes:

[0081] The images to be matched are scaled and normalized in size to obtain preprocessed images;

[0082] The preprocessed image is divided into N image blocks according to a preset block partitioning rule, and the image blocks are flattened to obtain image block vectors.

[0083] The image patch vector is linearly mapped to generate the image patch embedding vector, and a classification label embedding vector representing the whole image is added before the image patch embedding vector. The classification label embedding vector and the image patch embedding vector are concatenated to form the image embedding sequence.

[0084] The image embedding sequence is overlaid with position encoding. The image embedding sequence after overlay position encoding is input into the CLIP encoder in the improved BLIP-2 model. In the CLIP encoder, the image embedding sequence is extracted by a multi-head self-attention network and a feedforward network in sequence, and the corresponding visual feature sequence is output.

[0085] In this embodiment, step two specifically includes:

[0086] The visual feature sequence is input into the Q-Former module. M learnable query vectors are preset in the Q-Former module. The learnable query vectors are linearly transformed to obtain the query representation. The visual feature sequence is linearly transformed to obtain the key representation and value representation.

[0087] The query representation is input into a multi-head self-attention network to model the relationship between the learnable query vectors, thus obtaining the updated query representation.

[0088] The updated query representation, along with the key and value representations, is input into a multi-head cross-attention network. The correlation between the updated query representation and the visual feature sequence is calculated, and the image association representation corresponding to each learnable query vector is output.

[0089] The image association representation is subjected to a nonlinear transformation, and the image association representations are combined according to a preset method to generate an image semantic vector.

[0090] The text to be matched is divided into P text tokens according to a preset word segmentation rule, resulting in a text token sequence;

[0091] Each text token in the text token sequence is mapped to its corresponding word embedding vector, and a start token for identifying the start position of the text and an end token for identifying the end position of the text are added to the text token sequence to obtain an extended text token sequence containing the start and end tokens;

[0092] The word embedding vectors in the extended text tag sequence are superimposed with positional encoding to form a text embedding sequence with positional information;

[0093] The text embedding sequence with location information is input into the text encoder. The text embedding sequence is then processed by multi-head self-attention operation and feedforward operation to extract features and output the hidden representations corresponding to each text tag.

[0094] The hidden representations are aggregated according to a preset method to generate a text semantic vector that represents the overall semantics of the text to be matched.

[0095] In this embodiment, step three specifically includes:

[0096] Image semantic vectors and text semantic vectors are organized into preset batches to form image semantic vector sets and text semantic vector sets within the same batch;

[0097] In the image-text alignment module, the image semantic vector of the same image and the text semantic vector corresponding to the same image are set as image-text positive sample pairs;

[0098] Image semantic vectors and non-corresponding text semantic vectors within a batch are set as negative image-text sample pairs in the image-to-text direction;

[0099] Set the text semantic vector and the non-corresponding image semantic vector within the batch as a text-to-image negative sample pair;

[0100] The similarity between the semantic vector of each image in the batch and the semantic vector of each text in the batch is calculated to obtain the set of similarity between the image and the text.

[0101] Based on the similarity set from image to text and the preset temperature parameter, a normalization operation is performed to obtain the alignment probability distribution from image to text.

[0102] Based on the alignment probability distribution from image to text and the labeling information of positive image-text samples, calculate the alignment loss from image to text.

[0103] The similarity between each text semantic vector in the batch and each image semantic vector in the batch is calculated to obtain the set of similarity between text and image.

[0104] Based on the similarity set from text to image and the preset temperature parameter, a normalization operation is performed to obtain the alignment probability distribution from text to image.

[0105] Based on the alignment probability distribution of text to image direction and the labeling information of positive image-text sample pairs, calculate the alignment loss of text to image direction;

[0106] The alignment loss from image to text direction is summed with the alignment loss from text to image direction to obtain the InfoNCE alignment loss.

[0107] The trainable parameters in the Q-Former module and the text encoder are updated using gradients based on the InfoNCE alignment loss. The updated image semantic vector output by the Q-Former module is combined with the updated text semantic vector output by the text encoder to generate multimodal alignment features that represent the cross-modal alignment results between the image and the text.

[0108] The gradient update of the trainable parameters in the Q-Former module and the text encoder based on the InfoNCE alignment loss is as follows:

[0109] Using a backpropagation-based parameter optimization method, the image semantic vector and the text semantic vector are input into the image-text alignment module to calculate the bidirectional InfoNCE alignment loss, resulting in a scalar loss value representing the image-text alignment error.

[0110] Using the InfoNCE alignment loss as the objective function, the gradients of all trainable parameters involved in training in the Q-Former module and the text encoder are calculated using the backpropagation algorithm.

[0111] Based on gradients, the Adam optimizer is used to update each trainable parameter.

[0112] The above steps of loss calculation, gradient backpropagation and parameter update are iteratively executed according to the preset training rounds, so that the Q-Former module and the text encoder can gradually learn the cross-modal correspondence between images and text.

[0113] This invention proposes a large-scale model-based image-text matching method. By inputting image semantic vectors and text semantic vectors into an image-text alignment module, positive and negative image-text sample pairs are constructed, and a bidirectional InfoNCE alignment loss is calculated to quantify the image-text semantic matching error. Based on this, the backpropagation algorithm is used to differentiate the loss function, and the Adam optimizer is combined to update the gradients of the trainable parameters in the Q-Former module and the text encoder. This allows the model to continuously optimize the alignment capability between image and text representations in multiple iterations, ultimately obtaining more expressive multimodal alignment features and improving the accuracy of image-text matching.

[0114] In this embodiment, step four specifically includes:

[0115] The multimodal alignment features and the embedding vectors corresponding to the text tag sequence are concatenated along the feature dimension to obtain the joint input sequence;

[0116] Position encoding is superimposed on each embedding vector in the joint input sequence to obtain a joint embedding sequence with position information;

[0117] The joint embedding sequence is sequentially input into each layer of the decoder of the improved LLaMA;

[0118] In each decoder layer, before performing multi-head self-attention computation, the hidden state output by the previous decoder layer is obtained, and the hidden state is input into the learning attention path planning sub-network to generate the attention path planning matrix corresponding to the multi-head self-attention weight matrix of the current layer.

[0119] During the multi-head self-attention calculation process of the current layer, the multi-head self-attention weight matrix of the current layer is constrained element by element using the attention path planning matrix to obtain the multi-head self-attention weights constrained by the path planning.

[0120] The step of using the attention path planning matrix to constrain the multi-head self-attention weight matrix of the current layer element-wise involves multiplying the attention path planning matrix output by the learned attention path planning sub-network with the multi-head self-attention weight matrix of the current layer element-wise according to the corresponding positions. By performing element-wise multiplication, the attention weights are structurally constrained, so that the attention weights corresponding to the positions suppressed by the path planning matrix are reduced or masked, while the attention weights corresponding to the positions retained by the path planning matrix are preserved, thereby generating an attention weight matrix constrained by path planning.

[0121] The value representations corresponding to the key representations of the current layer decoder are weighted and summed based on the multi-head self-attention weights constrained by path planning to generate the updated hidden state of the current layer.

[0122] This invention introduces a learning attention path planning sub-network into an improved LLaMA decoder to structurally constrain the self-attention mechanism of each layer of the decoder. Specifically, multimodal alignment features are concatenated with text embedding vectors and input into the decoder. In each layer, the hidden state of the previous layer is used to generate an attention path planning matrix, which is then multiplied element-wise with the multi-head self-attention weight matrix of the current layer. This achieves constraint control over the attention connection relationship, thereby enhancing the model's ability to select attention regions in image-text semantic matching tasks and improving the structural expressiveness and matching discrimination ability of attention computation.

[0123] In this embodiment, the learning attention path planning subnetwork includes a cross-modal feature fusion layer, a content grouping and relation calculation layer, and a path planning matrix generation layer;

[0124] The cross-modal feature fusion layer performs linear mapping between the current layer's hidden state and the image semantic vector, and performs cross-modal fusion through concatenation and gating to obtain cross-modal fused features;

[0125] The content grouping and relation calculation layer performs nonlinear transformation on the cross-modal fusion features to generate content scores. Based on the content scores, a relation graph representing the relationship between multimodal alignment features and the hidden states corresponding to the text tag sequences is constructed through learnable soft clustering.

[0126] The step of constructing a relational graph based on content scores and learnable soft clustering to characterize the relationships between hidden states corresponding to sequence elements in the joint input sequence is as follows:

[0127] Q cluster center vectors are preset, each of which is a trainable parameter that is updated along with the network parameters during training. The content score is compared with each cluster center vector to obtain a similarity score set. The similarity score set is normalized to obtain a soft assignment weight vector. The soft assignment weights of each label are combined in a preset way to represent the degree of common assignment of any two labels on each cluster center and generate corresponding relationship values. The relationship values ​​of all labels are arranged in the label order to form a relationship graph, which is used to represent the strength of content relevance between any pair of labels.

[0128] The path planning matrix generation layer takes the relationship graph as input, linearly maps and activates the function to generate a path planning matrix with a range of element values. The path planning matrix is ​​then sparsified by sorting to obtain an attention path planning matrix used to constrain the multi-head self-attention connection relationship.

[0129] In a certain layer of the decoder in the improved LLaMA, the self-attention weight matrix is ​​generated and used to update the hidden state of the current layer through the following steps:

[0130] Step 1: Calculate the attention score matrix by matching the query representation of the current layer with the key representation to obtain the attention score matrix that characterizes the degree of association between the elements of each sequence.

[0131] Step 2: Generate the self-attention weight matrix and perform normalization processing on the attention score matrix so that the attention value corresponding to each query position meets the preset normalization requirements, thereby obtaining the self-attention weight matrix.

[0132] Step 3: Execute path planning constraints. Multiply the attention path planning matrix output by the attention path planning sub-network with the self-attention weight matrix element by element according to the corresponding positions. Apply structured constraints to the self-attention weight matrix to obtain the self-attention weight matrix constrained by path planning.

[0133] Step 4: Weighted synthesis of the current layer hidden state. Based on the self-attention weight matrix constrained by path planning, the value representations corresponding to the key representations of the current layer are weighted and summed to generate the updated hidden state of the current layer.

[0134] This invention optimizes the self-attention mechanism in an improved LLaMA model by designing a learning attention path planning sub-network. This sub-network consists of a cross-modal feature fusion layer, a content grouping and relation calculation layer, and a path planning matrix generation layer. The cross-modal feature fusion layer fuses the current layer's hidden state with the image semantic vector to generate cross-modal fused features. The content grouping and relation calculation layer generates content scores based on the fused features and constructs a relation graph between sequence elements using a soft clustering mechanism based on trainable cluster centers to represent their content relevance. The path planning matrix generation layer generates a sparse attention path planning matrix based on the relation graph, which constrains the connection relationships of the self-attention weights. Finally, in the decoder, this path planning matrix applies element-wise structured constraints to the self-attention weight matrix, and the constrained weights are used to perform a weighted summation of the value representations, thereby updating the current layer's hidden state and improving the model's ability to focus on key semantics in image-text matching and the accuracy of cross-modal modeling.

[0135] In this embodiment, the semantic consistency calibration mechanism specifically includes:

[0136] In the current layer decoder, the current layer hidden state after multi-head self-attention calculation and path planning constraints is obtained to obtain the current layer hidden state used for semantic consistency calibration.

[0137] Projection operations are performed on the image semantic vector to obtain a projection vector that is in the same feature space as the current layer hidden state;

[0138] The projection vector is fused with the current layer hidden state according to a preset method to form a fused representation for gating coefficient calculation;

[0139] Perform nonlinear and linear transformations on the fused representation to obtain gating coefficients corresponding to the hidden state dimension of the current layer;

[0140] The hidden state of the current layer is adjusted element by element based on the gating coefficient to obtain the calibrated hidden state after semantic consistency calibration;

[0141] The calibration hidden state is input into the feedforward network module, and linear and nonlinear transformations are performed on the calibration hidden state. The residual connection and normalization operation are then superimposed to generate the output of the current layer decoder.

[0142] This invention introduces a semantic consistency calibration mechanism into the decoder. By fusing the image semantic vector with the hidden state calculated by the current layer's self-attention mechanism and constrained by path planning, the semantic information is dynamically adjusted. Specifically, the image semantic vector is projected into the same feature space as the hidden state and fused with the hidden state to form a fused representation. Then, nonlinear and linear transformations are used to generate gating coefficients. Based on these gating coefficients, the hidden state is adjusted element-wise to obtain a calibrated hidden state with stronger semantic consistency. Finally, the calibrated hidden state is input into a feedforward network module for further processing, ensuring that the generated decoder output more accurately reflects the consistent semantic relationship between the image and text, thereby improving the accuracy and robustness of image-text matching.

[0143] In this embodiment, step six specifically includes:

[0144] Obtain the updated hidden state output from the top-level decoder of the improved LLaMA and use it as the hidden state input for the text-to-image matching regression.

[0145] The hidden state of the image-text matching regression is input into the output layer of the improved BLIP-2 model. In the output layer, linear and nonlinear transformations are performed on the hidden state of the image-text matching regression to obtain the intermediate matching vector used to represent the image-text matching relationship.

[0146] A regression operation is performed on the intermediate matching vector to generate an image-text matching score that characterizes the degree of matching between the image to be matched and the text to be matched.

[0147] Based on the comparison between the image-text matching score and the preset threshold, the corresponding image-text matching result is output.

[0148] Example 1:

[0149] To verify the effectiveness and advancement of this invention, the proposed image-text matching method based on the improved BLIP-2 and LLaMA decoding structures was applied to a real-world image-text semantic matching system. This system is primarily used to automatically identify the correspondence between images and natural language descriptions, and is widely applicable to scenarios such as smart libraries, intelligent surveillance, and cross-modal search engines. This paper uses image-text pairs from an open image-text dataset as test objects for experiments and compares the model performance with existing mainstream methods to comprehensively evaluate its capabilities.

[0150] In practical applications, 8,000 images were first collected, including categories such as people, animals, vehicles, natural landscapes, and daily life scenes, and 16,000 natural language descriptions were written to accompany them. The images came from diverse acquisition terminals, with a resolution of no less than 512×512, and all images underwent uniform size scaling and pixel normalization. The text descriptions were generated manually, with a length between 10 and 25 characters, ensuring coverage of the core semantic content of the images.

[0151] When applying the image-text matching method proposed in this invention, the image is first input into the CLIP encoder in the improved BLIP-2 model to extract the visual feature sequence. Then, the image features are transformed into image semantic vectors through the Q-Former module. The text semantic vector is obtained through the text encoder and then fed into the image-text alignment module to calculate the alignment error based on the InfoNCE loss function, achieving cross-modal semantic alignment optimization. Multimodal features are fed into the LLaMA decoder, which incorporates an attention path planning mechanism. A path planning sub-network is introduced at each layer to structurally constrain the attention weights, and a semantic consistency calibration mechanism is used to enhance the adaptability to image semantics, ultimately outputting the image-text matching score.

[0152] To comprehensively evaluate the matching performance of the method of this invention, the currently mainstream BLIP-2 standard model, ALBEF model, and CLIP model were selected as comparison baselines, and training and testing were conducted on the same dataset and hardware platform (using NVIDIA A100 GPU). Performance evaluation metrics included Top-1 accuracy, Top-5 accuracy, recall, F1 score, and average matching time. Specific experimental data are shown in Table 1:

[0153] Table 1. Comparison of the method of this invention with existing mainstream image-text matching methods

[0154] Method of the present invention 86.7 97.2 89.4 88.1 0.042 16.8 BLIP-2 Standard Edition 81.3 94.5 83.1 82.7 0.038 benchmark CLIP 78.5 91.4 79.8 79.2 0.033 - ALBEF 80.2 93.1 81.7 80.9 0.041 5.1

[0155] Table 1 shows that in the positive image-text matching task, the Top-1 accuracy of the method of this invention reaches 86.7%, significantly better than CLIP model's 78.5% and BLIP-2's 81.3%; the Top-5 accuracy reaches 97.2%, the recall rate is 89.4%, and the F1 score is 88.1%; in the image-text mismatch filtering task, the model of this invention can effectively reduce the mismatch rate, with an average reduction in misjudgments of 16.8%. Due to the introduction of a path planning matrix and gating fusion mechanism, the computational efficiency of this invention is only slightly higher than that of a conventional LLaMA structure, with an average processing time of 0.042 seconds per pair of image-text matching, meeting the real-time requirements of industrial applications.

[0156] The results show that this invention can effectively solve the problems of inaccurate semantic mapping between images and text and weak cross-modal correlation. By introducing path planning and semantic calibration mechanisms, it improves the model's ability to express image-text alignment relationships and enhances the system's intelligent semantic understanding capabilities. It demonstrates strong adaptability and scalability in large-scale image-text recognition and retrieval tasks.

[0157] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A text-image matching method based on a large model, characterized in that, Includes the following steps: Step 1: Obtain the image to be matched. Input the image to be matched into the CLIP encoder in the improved BLIP-2 model to obtain the visual feature sequence. The improved BLIP-2 model includes a CLIP encoder, a Q-Former module, a text encoder, an improved LLaMA, and an output layer, and introduces an image-text alignment module. Step 2: Input the visual feature sequence into the Q-Former module to obtain the image semantic vector; obtain the text to be matched, divide the text to be matched into a text tag sequence, input the text tag sequence into the text encoder to obtain the text semantic vector; Step 3: Input the image semantic vector and the text semantic vector into the image-text alignment module, calculate the alignment loss based on InfoNCE, and update the trainable parameters in the Q-Former module and the text encoder according to the alignment loss to obtain multimodal alignment features; Step 4: Concatenate the multimodal alignment features with the embedding vectors corresponding to the text tag sequence to form a joint input sequence. Input the joint input sequence into the decoder of each layer of the improved LLaMA. Before the self-attention calculation, generate the attention path planning matrix by learning the attention path planning sub-network, and use the attention path planning matrix to perform structured constraints on the attention weights of the self-attention weight matrix of the current layer. Step 5: After multi-head self-attention calculation, a semantic consistency calibration mechanism is introduced. The hidden state of the current layer is projected and fused according to the image semantic vector to generate gating coefficients. The hidden state of the current layer is adjusted according to the gating coefficients. The adjusted hidden state is input into the feedforward network module to generate the output of the current layer decoder. Step 6: Input the hidden state output by the top-level decoder into the output layer of the improved BLIP-2 model, perform regression calculation on the hidden state, and generate the image-text matching result; Step four specifically includes: The multimodal alignment features and the embedding vectors corresponding to the text tag sequence are concatenated along the feature dimension to obtain the joint input sequence; Position encoding is superimposed on each embedding vector in the joint input sequence to obtain a joint embedding sequence with position information; The joint embedding sequence is sequentially input into each layer of the decoder of the improved LLaMA; In each decoder layer, before performing multi-head self-attention computation, the hidden state output by the previous decoder layer is obtained, and the hidden state is input into the learning attention path planning sub-network to generate the attention path planning matrix corresponding to the multi-head self-attention weight matrix of the current layer. During the multi-head self-attention calculation process of the current layer, the multi-head self-attention weight matrix of the current layer is constrained element by element using the attention path planning matrix to obtain the multi-head self-attention weights constrained by the path planning. Based on the multi-head self-attention weights constrained by path planning, the value representations corresponding to the key representations of the current layer decoder are weighted and summed to generate the updated hidden state of the current layer. The learning attention path planning subnetwork includes a cross-modal feature fusion layer, a content grouping and relation calculation layer, and a path planning matrix generation layer; The cross-modal feature fusion layer performs linear mapping between the current layer's hidden state and the image semantic vector, and performs cross-modal fusion through concatenation and gating to obtain cross-modal fused features; The content grouping and relation calculation layer performs nonlinear transformation on the cross-modal fusion features to generate content scores. Based on the content scores, a relation graph is constructed through learnable soft clustering to characterize the relationship between the hidden states corresponding to the sequence elements in the joint input sequence. The path planning matrix generation layer takes the relationship graph as input, linearly maps and activates the function to generate a path planning matrix with a range of element values. The path planning matrix is ​​then sparsified by sorting to obtain an attention path planning matrix used to constrain the multi-head self-attention connection relationship.

2. The image-text matching method based on a large model according to claim 1, characterized in that, Step one specifically includes: The images to be matched are scaled and normalized in size to obtain preprocessed images; The preprocessed image is divided into N image blocks according to a preset block partitioning rule, and the image blocks are flattened to obtain image block vectors. The image patch vector is linearly mapped to generate the image patch embedding vector, and a classification label embedding vector representing the whole image is added before the image patch embedding vector. The classification label embedding vector and the image patch embedding vector are concatenated to form the image embedding sequence. The image embedding sequence is overlaid with position encoding. The image embedding sequence after overlay position encoding is input into the CLIP encoder in the improved BLIP-2 model. In the CLIP encoder, the image embedding sequence is extracted by a multi-head self-attention network and a feedforward network in sequence, and the corresponding visual feature sequence is output.

3. The image-text matching method based on a large model according to claim 1, characterized in that, Step two specifically includes: The visual feature sequence is input into the Q-Former module. M learnable query vectors are preset in the Q-Former module. The learnable query vectors are linearly transformed to obtain the query representation. The visual feature sequence is linearly transformed to obtain the key representation and value representation. The query representation is input into a multi-head self-attention network to model the relationship between the learnable query vectors, thus obtaining the updated query representation. The updated query representation, along with the key and value representations, is input into a multi-head cross-attention network. The correlation between the updated query representation and the visual feature sequence is calculated, and the image association representation corresponding to each learnable query vector is output. The image association representation is subjected to a nonlinear transformation, and the image association representations are combined according to a preset method to generate an image semantic vector. The text to be matched is divided into P text tokens according to a preset word segmentation rule, resulting in a text token sequence; Each text token in the text token sequence is mapped to its corresponding word embedding vector, and a start token for identifying the start position of the text and an end token for identifying the end position of the text are added to the text token sequence to obtain an extended text token sequence containing the start and end tokens; The word embedding vectors in the extended text tag sequence are superimposed with positional encoding to form a text embedding sequence with positional information; The text embedding sequence with location information is input into the text encoder. The text embedding sequence is then processed by multi-head self-attention operation and feedforward operation to extract features and output the hidden representations corresponding to each text tag. The hidden representations are aggregated according to a preset method to generate a text semantic vector that represents the overall semantics of the text to be matched.

4. The image-text matching method based on a large model according to claim 1, characterized in that, Step three specifically includes: Image semantic vectors and text semantic vectors are organized into preset batches to form image semantic vector sets and text semantic vector sets within the same batch; In the image-text alignment module, the image semantic vector of the same image and the text semantic vector corresponding to the same image are set as image-text positive sample pairs; Image semantic vectors and non-corresponding text semantic vectors within a batch are set as negative image-text sample pairs in the image-to-text direction; Set the text semantic vector and the non-corresponding image semantic vector within the batch as a text-to-image negative sample pair; The similarity between the semantic vector of each image in the batch and the semantic vector of each text in the batch is calculated to obtain the set of similarity between the image and the text. Based on the similarity set from image to text and the preset temperature parameter, a normalization operation is performed to obtain the alignment probability distribution from image to text. Based on the alignment probability distribution from image to text and the labeling information of positive image-text pairs, calculate the alignment loss from image to text. The similarity between each text semantic vector in the batch and each image semantic vector in the batch is calculated to obtain the set of similarity between text and image. Based on the similarity set from text to image and the preset temperature parameter, a normalization operation is performed to obtain the alignment probability distribution from text to image. Based on the alignment probability distribution of text to image direction and the labeling information of positive image-text sample pairs, calculate the alignment loss of text to image direction; The alignment loss from image to text direction is summed with the alignment loss from text to image direction to obtain the InfoNCE alignment loss. The trainable parameters in the Q-Former module and the text encoder are updated using gradients based on the InfoNCE alignment loss. The updated image semantic vector output by the Q-Former module is combined with the updated text semantic vector output by the text encoder to generate multimodal alignment features that represent the cross-modal alignment results between the image and the text.

5. The image-text matching method based on a large model according to claim 1, characterized in that, The semantic consistency calibration mechanism specifically includes: In the current layer decoder, the current layer hidden state after multi-head self-attention calculation and path planning constraints is obtained to obtain the current layer hidden state used for semantic consistency calibration. Projection operations are performed on the image semantic vector to obtain a projection vector that is in the same feature space as the current layer hidden state; The projection vector is fused with the current layer hidden state according to a preset method to form a fused representation for gating coefficient calculation; Perform nonlinear and linear transformations on the fused representation to obtain gating coefficients corresponding to the hidden state dimension of the current layer; The hidden state of the current layer is adjusted element by element based on the gating coefficient to obtain the calibrated hidden state after semantic consistency calibration; The calibration hidden state is input into the feedforward network module, and linear and nonlinear transformations are performed on the calibration hidden state. The residual connection and normalization operation are then superimposed to generate the output of the current layer decoder.

6. The image-text matching method based on a large model according to claim 1, characterized in that, Step six specifically includes: Obtain the updated hidden state output from the top-level decoder of the improved LLaMA and use it as the hidden state input for the text-to-image matching regression. The hidden state of the image-text matching regression is input into the output layer of the improved BLIP-2 model. In the output layer, linear and nonlinear transformations are performed on the hidden state of the image-text matching regression to obtain the intermediate matching vector used to represent the image-text matching relationship. A regression operation is performed on the intermediate matching vector to generate an image-text matching score that characterizes the degree of matching between the image to be matched and the text to be matched. Based on the comparison between the image-text matching score and the preset threshold, the corresponding image-text matching result is output.

Citation Information

Patent Citations

  • Attention head-based large language model function partition detection method and system

    CN120317284A

  • Single cell transcriptome data and text description conjoint analysis method based on multi-modal language model

    CN120452543A