Cross-modal image-text retrieval processing method and system

Through the cross-modal graphic search processing method, the cross-modal attention mechanism is used to fuse the image and text features, and the cross-modal graphic search model is used to match similarity, which solves the problem of low accuracy in the existing technology, and achieves more efficient and accurate cross-modal retrieval.

CN119988664AInactive Publication Date: 2025-05-13LONGSHINE TECH

Patent Information

Application Number
CN202510460264.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The graphic and text search processing method in the prior art has the defect of low search accuracy and poor results, and has failed to fully utilize the complementarity between the image and the text, resulting in failure to capture sufficient context information in the initial search stage.

Method used

Through the cross-modal graphic and text search processing method, the user query text is obtained and encoded to generate query text feature vectors; the image and text features are modally fused based on the cross-modal attention mechanism to generate multimodal embedding representations; the cross-modal graphic and text search model is used to match similarity with the multimodal embedding representation in the external knowledge base, return relevant results, and perform image questions and answers with text assistance when necessary to improve the accuracy of the search results.

Benefits of technology

By better capturing semantic relationships between images and text, improve the accuracy and efficiency of graphic and text retrieval, reduce interference with low quality or irrelevant results, and improve the accuracy and interpretability of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988664A_ABST
    Figure CN119988664A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal image-text retrieval processing method and a cross-modal image-text retrieval processing system, which are applied to the field of information retrieval, and the method comprises the following steps: obtaining a query text input by a user in an image-text retrieval process; encoding the query text through a text encoder to generate a query text feature vector; through a cross-modal image-text retrieval model, similarity matching is carried out based on query text feature vectors and multi-modal embedding representations stored in an external knowledge base, related results corresponding to the multi-modal embedding representations larger than a matching threshold value are returned, and the multi-modal embedding representations are used for representing joint features of images and texts; when the related result comprises the image and the text at the same time, the related result and the query text are input into a preset multi-modal large model, image question answering with text assistance is carried out, and a retrieval result output by the multi-modal large model is obtained; according to the method, the semantic association between the image and the text can be better captured, so that the accuracy of image-text retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval technology, and in particular to a cross-modal image and text retrieval processing method and system. Background Art

[0002] In the past, information retrieval technology mainly focused on a single modality, such as keyword-based text retrieval or image retrieval. Cross-modal image-text retrieval allows users to use data in one modality (image or text) as a query and return the results with the highest similarity in another modality.

[0003] Existing modality fusion methods fail to fully utilize the complementarity between image and text information when processing them. This deficiency may result in a failure to capture sufficient contextual information in the initial retrieval stage, thus affecting the accuracy and effectiveness of retrieval.

[0004] It can be seen from this that the image and text retrieval processing method in the related art has technical problems such as low retrieval accuracy and poor results. Summary of the invention

[0005] The present invention provides a cross-modal image-text retrieval processing method and system to solve the defects of the image-text retrieval processing method in the prior art, such as low retrieval accuracy and poor effect, so as to better capture the semantic association between images and texts, thereby improving the accuracy of image-text retrieval.

[0006] The present invention provides a cross-modal image-text retrieval processing method, comprising the following steps.

[0007] The query text input by the user in the image-text retrieval process is obtained; the query text is encoded by a text encoder to generate a query text feature vector; based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base, similarity matching is performed through a cross-modal image-text retrieval model, and a target number of related results corresponding to the multimodal embedding representation greater than a matching threshold are returned, wherein the multimodal embedding representation is used to represent the joint features of the image and the text, and the related results include at least one of the following: image and text; when the related results include both the image and the text, the related results and the query text are input into a preset multimodal large model, and image question and answer with text assistance is performed to obtain the retrieval results output by the multimodal large model.

[0008] According to a cross-modal image and text retrieval processing method provided by the present invention, before the similarity calculation is performed based on the query text feature vector and the text features and image features stored in the external knowledge base, the method also includes: performing image and text segmentation on the input document to obtain text data and image data; encoding the image data through an image encoder to obtain an input image feature vector; encoding the text data through a text encoder to obtain an input text feature vector; based on a cross-modal attention mechanism, performing modal fusion on the input text feature vector and the input image feature vector to obtain a multimodal embedding representation; and constructing an external knowledge base based on the multimodal embedding representation.

[0009] According to a cross-modal image and text retrieval processing method provided by the present invention, the image encoder is a Transformer-based model, and the image data is encoded by the image encoder to obtain an input image feature vector, including: dividing the image data into multiple image blocks of a fixed size; mapping each image block in the multiple image blocks to an embedding space of a fixed length to obtain an embedding vector of each image block; position encoding each image block to obtain the relative position of each image block; and determining the input image feature vector corresponding to the image data based on the embedding vector of each image block and the relative position of each image block according to a self-attention mechanism.

[0010] According to a cross-modal image-text retrieval processing method provided by the present invention, based on the cross-modal attention mechanism, the input text feature vector and the input image feature vector are modally fused to obtain a multimodal embedding representation, including: determining a first attention weight between the i-th input image feature vector and the j-th input text feature vector; determining a second attention weight between the j-th input text feature vector and the i-th input image feature vector; determining a first fusion result of the input image feature vector on the input text feature vector based on the first attention weight; determining a second fusion result of the input text feature vector on the input image feature vector based on the second attention weight; and combining the first fusion result with the second fusion result to obtain a multimodal embedding representation.

[0011] According to a cross-modal image-text retrieval processing method provided by the present invention, before the cross-modal image-text retrieval model is used to perform similarity calculation based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base, and the relevant results corresponding to the multimodal embedding representation with the highest similarity of the target number are returned, the method also includes: obtaining an image-text sample pair training set, wherein the image-text sample pair training set includes positive sample pairs and negative sample pairs; determining a first similarity between the image features and the text features of the positive sample pair and a second similarity between the image features and the text features of the negative sample pair; determining a contrast loss function based on the first similarity and the second similarity; and performing contrast loss training on the cross-modal image-text retrieval model based on the contrast loss function, wherein the contrast loss function is: ; in, represents the contrast loss function, represents the total number of sample pairs of the image-text sample pair training set, represents the index of the positive sample pair, Indicates The first similarity of positive sample pairs, represents the index of the negative sample pair, Indicates selecting the maximum value. Indicates The second similarity of the negative sample pairs, Indicates a preset first similarity threshold.

[0012] According to a cross-modal image-text retrieval processing method provided by the present invention, after performing contrast loss training on the cross-modal image-text retrieval model based on the contrast loss function, the method further includes: obtaining a query sample pair; determining a difficult positive sample pair and a difficult negative sample pair corresponding to the query sample pair, wherein the distance between the difficult positive sample pair and the query sample pair in the feature space is greater than a distance threshold, and the distance between the difficult negative sample pair and the query sample pair in the feature space is less than a distance threshold; constructing image triples and text triples based on the query sample pair, the difficult positive sample pair and the difficult negative sample pair; determining the query image features, positive sample image features and negative sample image features of the image triples; determining the query text features, positive sample text features and negative sample text features of the text triples; determining a triple loss function based on the query image features, the positive sample image features, the negative sample image features, the query text features, the positive sample text features and the negative sample text features; performing triple loss training on the cross-modal image-text retrieval model based on the triple loss function, wherein the triple loss function is: ; in, represents the triple loss function, Indicates selecting the maximum value. represents the distance function, represents the query image feature, represents the positive sample image feature, represents the negative sample image feature, Indicates the preset first boundary value, represents the query text feature, represents the positive sample text feature, represents the negative sample text feature, Indicates the preset second boundary value.

[0013] The present invention also provides a cross-modal image-text retrieval processing system, comprising the following modules: An acquisition module is used to acquire a query text input by a user during a graphic and text retrieval process; an encoding module is used to encode the query text through a text encoder to generate a query text feature vector; a similarity module is used to perform similarity matching based on the query text feature vector and a multimodal embedding representation stored in an external knowledge base through a cross-modal graphic and text retrieval model, and return a target number of related results corresponding to multimodal embedding representations greater than a matching threshold, wherein the multimodal embedding representation is used to represent the joint features of an image and text, and the related results include at least one of the following: image and text; a retrieval result module is used to input the related results and the query text into a preset multimodal large model when the related results include both an image and text, perform image question and answer with text assistance, and obtain a retrieval result output by the multimodal large model.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the cross-modal image and text retrieval processing method as described above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the cross-modal image-text retrieval processing methods described above.

[0016] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned cross-modal image-text retrieval processing methods.

[0017] The cross-modal image-text retrieval processing method and system provided by the present invention converts the query text input by the user into a high-dimensional feature vector through a text encoder, which can capture the semantic information of the text and provide a basis for subsequent cross-modal matching; the multimodal embedding representation stored in the external knowledge base (such as the joint features of images and texts) can effectively characterize the semantic relationship of data in different modalities and realize cross-modal similarity matching; through the cross-modal image-text retrieval model, the query text feature vector is matched with the multimodal embedding representation, which can quickly screen out results related to the semantics of the query text and improve the retrieval efficiency; by setting a matching threshold, it is ensured that the semantic relevance of the returned results to the query text is high enough to avoid low-quality or irrelevant results from interfering with the user; when the relevant results contain both images and texts, the multimodal large model is used to perform image question answering with text-assisted information to improve the accuracy and interpretability of the retrieval results, thereby realizing efficient and accurate cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a schematic diagram of the multimodal retrieval enhancement generation task processing flow of the related technology provided by the present invention.

[0020] Figure 2 It is a flowchart of the cross-modal image and text retrieval processing method provided by the present invention.

[0021] Figure 3 It is a schematic diagram of the overall flow of the cross-modal image and text retrieval processing method provided by the present invention.

[0022] Figure 4 It is a module schematic diagram of the cross-modal image and text retrieval processing system provided by the present invention.

[0023] Figure 5 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] In the past, information retrieval technology mainly focused on a single modality, such as keyword-based text retrieval or image retrieval. Cross-modal image-text retrieval allows users to use data in one modality (image or text) as a query and return the results with the highest similarity in another modality. However, this process faces the "heterogeneous gap" problem, that is, it is necessary to establish effective feature processing and mapping between modal data with different distributions and feature representations to accurately calculate the similarity of cross-modal data.

[0026] In recent years, the rapid development of deep learning has significantly promoted the progress of cross-modal image and text retrieval. By extracting and aligning features through deep networks, we can build a retrieval model that far exceeds the performance of traditional methods.

[0027] Due to the high efficiency of the Transformer model in natural language tasks, many scholars use transfer learning to use the Transformer as the backbone network for image and text retrieval tasks. According to its network structure, it can be divided into single-channel models, dual-channel models, and pre-trained models.

[0028] Single-Stream Architecture: Text and images are regarded as two different sets of data streams, which are transmitted into different channels respectively, and then an interaction module is used for data interaction, fusion and similarity calculation. However, its disadvantage is that it may lead to information loss and low computational efficiency, especially in the fusion stage, it is difficult to fully capture the semantic relationship between the two.

[0029] Double-Stream Architecture: Text and images are considered as the same data stream and are input together, and data interaction and fusion are performed directly at the beginning of the task. Compared with the single-channel model, the parallel structure of the dual-channel model has the advantage of high speed and can specifically extract features of different modalities. However, the parallel structure also has certain limitations, such as the fusion strategy may affect the retrieval effect and consume a lot of computing resources.

[0030] Pre-training Architecture: Text and images are generally considered as the same data stream and input together. Unlike the dual-channel model, the pre-training model is a large-scale model for multiple tasks, with a deep network structure and large parameters. It needs to be trained with a large amount of data and can handle multiple image-text related tasks, such as image-text retrieval, title generation, image understanding, etc. However, its high training cost, high risk of overfitting, and complex network structure make debugging and optimization more difficult.

[0031] After training with massive amounts of data, large multimodal models have demonstrated powerful capabilities in performing common natural language processing tasks. However, hallucination phenomena and challenges in updating parameter knowledge limit their practical applications in specific fields. The Retrieval-augmented Generation (RAG) method alleviates this problem by introducing a knowledge retriever that can access a customized external knowledge base, thereby providing the large model with the information needed to generate output.

[0032] However, current RAG systems are mainly based on text and fail to fully utilize visual information such as layout and images, which play a key role in real-world multimodal documents. The success of multimodal large models makes it possible to implement native multimodal RAG algorithms because these models can directly understand images without converting them into text, thereby being able to extract embeddings and perform RAG processing.

[0033] refer to Figure 1 , Figure 1 It is a schematic diagram of the multimodal retrieval enhancement generation task processing flow of the related technology provided by the present invention, which includes: document, page image, VisRAG-Ret, image database, retrieval data, Query, VisRAG-Ret, retrieved page, VisRAG-Gen and generated answer.

[0034] Currently, the multimodal retrieval enhancement generation task can effectively use vision-language models (VLM) to generate high-quality contextual embeddings from images of document pages. Figure 1 As shown in the figure, specifically, the VisRAG system consists of a VLM-based retriever VisRAG-Ret and a generator VisRAG-Gen. VisRAG-Ret inherits the dual encoder architecture of the traditional text-based dense retriever, mapping queries and documents to embedding spaces, but its innovation lies in directly utilizing the image of the document instead of relying on the extracted text content. This approach not only improves the efficiency of information utilization, but also ensures that more contextual information can be captured when processing complex documents. The generation of embeddings is obtained by weighted average pooling of the final hidden states of the input text and visual tags, which ensures the effective integration of different types of information.

[0035] After retrieving the top-k document images, VisRAG further processes these images to generate the final answer. Throughout the process, VisRAG retains all information in its original visual format, a feature that significantly reduces the risk of information loss or distortion that may occur in traditional RAG pipelines. In this way, VisRAG not only improves the accuracy of generated answers, but also enhances the model's ability to understand visual information, making it perform better in multimodal tasks.

[0036] However, existing modal fusion methods fail to fully utilize the complementarity between images and text when processing them. This deficiency may result in a failure to capture sufficient contextual information in the initial retrieval stage, thus affecting the accuracy and effectiveness of retrieval. By performing modal fusion on images and text, multimodal information can be better integrated, the overall performance of the retrieval system can be improved, and the relationship between images and text can be closer.

[0037] The current model only uses cross entropy loss, and this single loss function may not be able to fully capture the complexity and diversity of multimodal data. By introducing more complex loss functions, such as triple loss function and contrast loss function, the embedding space of the model can be better optimized, making similar queries and documents closer in the embedding space.

[0038] Currently, the training of multimodal RAG mainly relies on English data, which limits its adaptability and effectiveness in processing Chinese text and visual information. Therefore, the lack of training data for the Chinese context makes the system face many severe challenges in understanding and generating Chinese content.

[0039] The purpose of the present invention is to provide a cross-modal image and text retrieval processing method and system. In view of the shortcomings of the prior art, a modal fusion method and an improved training method are proposed. The present invention proposes a new modal fusion method, which effectively fuses the feature representations of images and texts and combines relevant text information with image features, thereby improving the accuracy of image and text retrieval. In addition, by improving the training method, the present invention combines (cross-modal) contrast loss and triple loss to better capture the semantic association between images and texts. Specifically, the cross-modal contrast loss is used to maximize the similarity between matching image and text pairs, while minimizing the similarity between unmatched pairs; the triple loss further improves the retrieval effect by modifying the data set format.

[0040] Optionally, the cross-modal image and text retrieval processing method of the embodiment of the present application can be executed by a server, or by a terminal device, or jointly by a server and a terminal device, taking the example of the cross-modal image and text retrieval processing method in the present embodiment being executed by a server.

[0041] Figure 2 is a flow chart of the cross-modal image and text retrieval processing method provided by the present invention, such as Figure 2 As shown, the method includes the following: Step 201, obtaining the query text input by the user during the image and text retrieval process.

[0042] In the embodiment of the present invention, when performing image and text retrieval, the user first inputs a question (query text), which is encoded by a text encoder to generate a corresponding query text feature vector.

[0043] Use front-end technologies (such as JavaScript, React, Vue, etc.) to monitor user input events (such as keyboard input or voice input) in the input box.

[0044] For example, during the user input process, the query text content is captured and stored in real time, and a simple check is performed on the text entered by the user, such as checking whether it is empty or contains illegal characters, and corresponding prompts are given.

[0045] Before inputting the query text into the text encoder, a preprocessing step is usually required, for example, removing extra spaces, line breaks, punctuation marks, etc.; using a word segmentation tool (such as Jieba, HanLP, etc.) to segment the query text into word sequences.

[0046] In some embodiments, the preprocessed query text is stored in a temporary variable or cache for subsequent processing, and the query text is passed to the backend server via an API or a network request (such as an HTTP POST request) for subsequent text encoding and retrieval operations.

[0047] Step 202: Encode the query text by using a text encoder to generate a query text feature vector.

[0048] In an embodiment of the present invention, the user's question is processed through a text encoder to obtain a text feature vector, wherein the text encoder may be a pre-trained language model, a traditional word vector model or a lightweight model.

[0049] Step 203, through the cross-modal image-text retrieval model, similarity matching is performed based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base, and relevant results corresponding to the multimodal embedding representation of the target number greater than the matching threshold are returned.

[0050] Among them, the multimodal embedding representation is used to represent the joint features of the image and the text, and the relevant results include at least one of the following: image and text.

[0051] In an embodiment of the present invention, a similarity calculation is performed between a query text feature vector corresponding to a query text input by a user and a multimodal embedding representation stored in an external knowledge base; and based on the result obtained by the similarity calculation, relevant results corresponding to the first k multimodal embedding representations are returned. Commonly used similarity measurement methods include cosine similarity, Euclidean distance, etc. In this way, the content most similar to the user's question can be found, whether it is text or image.

[0052] In an embodiment of the present invention, a cross-modal image-text retrieval model is used to match a query text feature vector with a multimodal embedding representation (such as joint features of images and text) in an external knowledge base, and return a target number of relevant results.

[0053] In some embodiments, the matching threshold and target number can be set according to the actual retrieval task requirements, and the multimodal embedding representation corresponding to the target number whose similarity is greater than the matching threshold is returned to ensure that the returned relevant results have a sufficiently high semantic relevance to the query text.

[0054] For example, all multimodal embedding representations are sorted from high to low by similarity, and relevant results (including images or texts) corresponding to the first N multimodal embedding representations are intercepted and returned.

[0055] Step 204, when the relevant results include both images and texts, the relevant results and the query text are input into a preset multimodal macro model, and image question answering with text assistance is performed to obtain the retrieval results output by the multimodal macro model.

[0056] In the embodiment of the present invention, when only text is returned as relevant result, the query text entered by the user and the retrieved text are simultaneously input into the large language model to obtain the corresponding answer. In this process, the large language model combines the query text of the user and the retrieved text information to generate a coherent and contextual answer.

[0057] When the only relevant results returned are images, the query text entered by the user and the retrieved images are simultaneously fed into the visual language model to obtain the corresponding answer. The visual language model combines the image information and the text information to generate a coherent and contextual answer.

[0058] When the returned relevant results include both images and texts, the query text input by the user and the retrieved images and texts are input into a preset multimodal macro model to obtain corresponding answers.

[0059] In an embodiment of the present invention, the input of the multimodal large model includes: query text and relevant results, wherein the query text is the query text entered by the user during the image and text retrieval process, and the relevant results include image data obtained from an external knowledge base, as well as text descriptions or contextual information related to the image.

[0060] In some embodiments, query text and related results (images and text) are obtained from a cross-modal image-text retrieval model; a visual encoder is used to convert the image into an image feature vector, and a text encoder is used to convert the query text and related text into a text feature vector; the image feature vector and the text feature vector are input into a multimodal large model for joint reasoning; the multimodal large model outputs an image question answer with text auxiliary information and returns it to the user. In this way, the multimodal large model can combine text information and image information to generate richer and more accurate answers.

[0061] Through the above steps of the embodiment of the present invention, the query text input by the user is converted into a high-dimensional feature vector through a text encoder, which can capture the semantic information of the text and provide a basis for subsequent cross-modal matching; the multimodal embedding representation stored in the external knowledge base (such as the joint features of images and texts) can effectively characterize the semantic relationship of data in different modalities and realize cross-modal similarity matching; through the cross-modal image-text retrieval model, the query text feature vector is matched with the multimodal embedding representation, which can quickly screen out results related to the semantics of the query text and improve the retrieval efficiency; by setting a matching threshold, it is ensured that the semantic relevance of the returned results to the query text is high enough to avoid low-quality or irrelevant results from interfering with users; when the relevant results contain both images and texts, the multimodal large model is used to perform image question answering with text-assisted information to improve the accuracy and interpretability of the retrieval results, thereby realizing efficient and accurate cross-modal retrieval.

[0062] According to a cross-modal image-text retrieval processing method provided by the present invention, before similarity matching is performed based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base, the method further includes: Perform image-text segmentation on the input document to obtain text data and image data; Encode the image data through an image encoder to obtain an input image feature vector; Encode the text data through a text encoder to obtain an input text feature vector; Based on the cross-modal attention mechanism, the input text feature vector and the input image feature vector are modally fused to obtain a multimodal embedding representation; Build external knowledge base based on multimodal embedding representation.

[0063] In the embodiment of the present invention, first, the input document is parsed and divided into two parts: text data and image data. The text data part includes all text content, and the image data part includes all image content. In this way, data of different modalities can be processed separately to ensure that their respective characteristics are optimally processed.

[0064] Next, the text block (text data) will be encoded by the text encoder to generate an input text feature vector. The text encoder can be a pre-trained language model. At the same time, the image block (image data) will be encoded by the image encoder to generate an input image feature vector. Then, the input text feature vector and the input image feature vector are input into the modality fusion module.

[0065] Finally, the results after modal fusion are stored in an external knowledge base to provide support for subsequent retrieval tasks.

[0066] The modal fusion module is used to fuse the input text feature vector and the input image feature vector into a unified multimodal embedding representation. Common fusion methods include: Concatenation: concatenate the feature vectors of text and image into a longer vector. Weighted summation: weighted summation of the feature vectors of text and image. Attention mechanism: use the attention mechanism to dynamically fuse the features of text and image. Multimodal embedding representation can capture the semantic information of text and image at the same time.

[0067] The multimodal embedding representation is stored in an external knowledge base together with the original text, image and its associated information. The external knowledge base can be a relational database (such as MySQL), a NoSQL database (such as MongoDB) or a vector database (such as FAISS).

[0068] refer to Figure 3 , Figure 3 It is a schematic diagram of the overall process of the cross-modal image and text retrieval processing method provided by the present invention, which includes: document, chunk segmentation, text, text encoder, picture, image encoder, query, similarity calculation, data storage library, Top k recall, returning text / picture according to the result, and inputting a large language model when there is only text; performing picture question and answer with auxiliary information when both pictures and text exist; and inputting a visual language model when there is only pictures. The specific process can be referred to the above embodiment and will not be repeated here.

[0069] Through the embodiment of the present invention, the steps of building an external knowledge base include segmenting the document, encoding text and image separately, fusing multimodal features and storing them in the knowledge base. This process can convert multimodal data into a unified embedding representation, providing basic support for cross-modal image and text retrieval.

[0070] According to a cross-modal image and text retrieval processing method provided by the present invention, the image encoder is a Transformer-based model, and the image encoder is used to encode the image data to obtain an input image feature vector, including: Dividing the image data into a plurality of image blocks of fixed size; Mapping each image block of the multiple image blocks to an embedding space of fixed length to obtain an embedding vector of each image block; Perform position encoding on each image block to obtain the relative position of each image block; According to the self-attention mechanism, the input image feature vector corresponding to the image data is determined based on the embedding vector of each image block and the relative position of each image block.

[0071] In the embodiment of the present invention, for the image feature extraction part, a Transformer-based model (such as Vision Transformer, ViT) is used to extract the features of the input image I (i.e., image data). This process includes the following steps: Image Blocking: ViT divides the image data into P fixed-size image patches, each of which is flattened and mapped to an embedding space of fixed length. Assume that the embedding vector of each image patch is , where i represents the index of the image block (from 1 to P).

[0072] Position encoding: Since the Transformer structure does not have the ability to process position information, it is necessary to add position encoding to each image block to ensure that the model can understand the relative positions between image blocks. Position encoding can help the model capture spatial information and avoid losing important spatial structures.

[0073] Self-attention mechanism: Through multiple self-attention layers, the model is able to learn the relationship between image patches and capture global context information. Transformer's self-attention mechanism allows each image patch to pay attention to all other patches, thereby better understanding the image content.

[0074] Finally, the input image feature vector (i.e., the embedded representation of the image) It is given by the following formula:

[0075] in, represents the input image feature vector, Represents a Transformer-based model, represents image data, Represents the embedding vector of each image block (i ranges from 1 to P).

[0076] Through the embodiments of the present invention, the Transformer-based model not only extracts the features of each image block, but also enhances the feature representation of the entire image through the self-attention mechanism, ensuring that the model can process complex patterns and contexts in the image.

[0077] In some embodiments, encoding the text data by a text encoder to obtain an input text feature vector specifically includes the following steps: For the input text T (text data), a pre-trained language model is used to extract its input text feature vector (embedded representation) The process includes: Tokenization: Divide text data into words or subwords and convert them into corresponding word embedding vectors , where j represents the index of the word (from 1 to N, where N is the number of words in the text data).

[0078] Context modeling: Using the self-attention mechanism, the pre-trained language model can capture the contextual relationship between words and generate a contextual embedding representation of each word. The pre-trained language model (Bidirectional Encoder Representations from Transformers, BERT) uses bidirectional encoding to ensure that the representation of each word takes into account its left and right contexts, thereby generating a more semantically meaningful embedding.

[0079] The final representation of the input text feature vector is given by the following formula:

[0080] in, represents the input text feature vector, represents the pre-trained language model, Represents text data, Represents the word embedding vector of each word (j ranges from 1 to N).

[0081] Through the embodiments of the present invention, the pre-trained language model can capture the semantic information in the text and the relationship between words, laying the foundation for subsequent modal fusion.

[0082] In some embodiments, weighted average fusion is performed on the first fusion result and the second fusion result.

[0083] The goal of modal fusion is to effectively combine image and text features to construct a unified multimodal feature representation. The weighted average method is used for fusion, and the weights are set. and (satisfy + =1, and ), the fused feature representation It is expressed as:

[0084] in, represents the fused feature representation, represents the first weight, represents the second weight, represents the input image feature vector, Represents the input text feature vector.

[0085] The choice of weights not only depends on experimental results, but can also be adjusted dynamically through feedback from model training. For example, in some applications, it may be found that image features contribute more to retrieval, so the weights can be increased. , in order to enhance the impact of image information.

[0086] According to a cross-modal image-text retrieval processing method provided by the present invention, based on a cross-modal attention mechanism, a modal fusion is performed on an input text feature vector and an input image feature vector to obtain a multi-modal embedding representation, including: Determine a first attention weight between the i-th input image feature vector and the j-th input text feature vector; Determine a second attention weight between the j-th input text feature vector and the i-th input image feature vector; Determine a first fusion result of the input image feature vector on the input text feature vector based on the first attention weight; Determine a second fusion result of the input text feature vector on the input image feature vector based on the second attention weight; The first fusion result and the second fusion result are combined to obtain a multimodal embedding representation.

[0087] It should be noted that in terms of feature extraction, the existing feature extraction methods fail to fully utilize the complementarity between image information and text information when processing them. In order to solve this problem, the modal fusion proposed in the present invention mainly includes the following key steps, which aim to optimize the fusion process of image and text information, thereby improving the performance of cross-modal image-text retrieval.

[0088] In the embodiment of the present invention, a cross-modal attention mechanism is introduced to calculate the attention weight A between the input image feature vector and the input text feature vector. The core of this step is to enhance the information exchange between modalities through the attention mechanism.

[0089] For the input image feature vector (i-th image block) and the input text feature vector (jth word), calculate the attention weight :

[0090] in, represents the first attention weight between the i-th input image feature vector and the j-th input text feature vector (i.e., the image-to-text attention weight), represents the second attention weight between the j-th input text feature vector and the i-th input image feature vector (i.e., the attention weight from text to image), represents the i-th input image feature vector, represents the jth input text feature vector, represents the similarity score, represents the exponential function used to map the similarity score to the positive range, ensuring the non-negativity of the attention weight. Used to exponentiate and sum the similarity scores between the i-th input image feature vector and all input text feature vectors (for normalization); Used to exponentiate and sum (for normalization) the similarity scores between the ith input text feature vector and all input image feature vectors.

[0091] in, Dot product or cosine similarity is usually used to calculate similarity. For example:

[0092] This self-attention mechanism allows the information of each modality to be weighted by weights, highlighting important information and suppressing unimportant information, thereby achieving effective information fusion.

[0093] Determine the first fusion result of the input image feature vector based on the first attention weight on the input text feature vector, which can be expressed by the following formula:

[0094] in, represents the first fusion result, represents the first attention weight between the i-th input image feature vector and the j-th input text feature vector, represents the jth input text feature vector, i represents the index of the input image feature vector, and the total number is P. Represents the index of the input text feature vector, the total number is N.

[0095] This formula means paying attention to text features through image features to obtain a text feature representation that integrates image attention information (i.e., the first fusion result).

[0096] Determine the second fusion result of the input text feature vector based on the second attention weight on the input image feature vector, which can be expressed by the following formula:

[0097] in, represents the second fusion result, represents the second attention weight between the j-th input text feature vector and the i-th input image feature vector, represents the i-th input image feature vector, i represents the index of the input image feature vector, and the total number is P. Represents the index of the input text feature vector, the total number is N.

[0098] This formula means paying attention to image features through text features to obtain an image feature representation that integrates text attention information (i.e., the second fusion result).

[0099] The first fusion result and the second fusion result are combined to obtain a multimodal embedding representation, which can be expressed by the following formula:

[0100] in, represents the multimodal embedding representation, represents the first fusion result, Represents the second fusion result.

[0101] In some embodiments, the resulting multimodal embedding representation (i.e., joint embedding) is Combining the feature information of images and texts can better capture the correlation information between the two. This joint representation is the basis for effective retrieval. In the retrieval process, similarity calculations (such as cosine similarity or Euclidean distance) are used to evaluate the relevance of the query image or text with other modalities in the database to achieve efficient retrieval:

[0102] in, represents the similarity calculation function, represents the multimodal embedding representation, Represents the query text feature vector.

[0103] Through the embodiments of the present invention, this combination method ensures that the semantic association between image and text features is enhanced, thereby improving the effect of modal fusion. The introduction of the cross-modal attention mechanism makes the feature representation not only a simple addition, but also generates a richer representation by strengthening complementary information.

[0104] The two-stage training method proposed in the present invention aims to improve the performance of the cross-modal image-text retrieval model, and specifically includes the following two stages: contrast loss training and triple loss training.

[0105] According to a cross-modal image-text retrieval processing method provided by the present invention, before performing similarity calculation based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base through the cross-modal image-text retrieval model and returning the relevant results corresponding to the multimodal embedding representation with the highest similarity of the target number, the method further includes: Obtaining an image-text sample pair training set, wherein the image-text sample pair training set includes positive sample pairs and negative sample pairs; Determine a first similarity between the image feature and the text feature of the positive sample pair and a second similarity between the image feature and the text feature of the negative sample pair; Determining a contrast loss function based on the first similarity and the second similarity; The cross-modal image-text retrieval model is trained with contrast loss based on the contrast loss function, where the contrast loss function is:

[0106] in, represents the contrast loss function, Represents the total number of sample pairs in the training set of image-text samples. represents the index of the positive sample pair, Indicates The first similarity of positive sample pairs, represents the index of the negative sample pair, Indicates selecting the maximum value. Indicates The second similarity of negative sample pairs, Indicates a preset first similarity threshold.

[0107] In the embodiment of the present invention, the image-text sample pair training set includes multiple image-text sample pairs. ,in, is the u-th image sample, is the text sample to match, Representation and Related description.

[0108] The image-text sample pair training set includes positive sample pairs and negative sample pairs: Positive sample pairs: contain image samples and text samples that match each other, i.e. .

[0109] Negative sample pairs: contain mismatched image samples and text samples, i.e. ,in is with Unrelated text samples.

[0110] Use the Transformer-based model to extract image features for each image sample: ,in, represents the input image feature vector, Represents a Transformer-based model, Represents image samples (i.e. image data).

[0111] Use the pre-trained language model to extract text features for each text sample: ,in, represents the input text feature vector, represents the pre-trained language model, Represents a text sample (i.e., text data).

[0112] In this way, the image sample and text sample of each image-text sample pair are mapped into a high-dimensional feature space.

[0113] Defining the similarity function , cosine similarity is usually used:

[0114] in, represents the similarity function, represents the first input, Represents the second input.

[0115] For each positive sample pair , calculate their similarity:

[0116] in, represents the first similarity of the u-th positive sample pair (between image features and text features), represents the similarity function, represents the image features of the u-th image sample, Represents the text features of the matching text sample.

[0117] For each negative sample pair , calculate their similarity:

[0118] in, represents the second similarity of the vth negative sample pair (between image features and text features), represents the similarity function, represents the image features of the u-th image sample, Represents text features of unrelated text samples.

[0119] Based on the first similarity and the second similarity, determine the contrast loss function The relevant formulas can be found above and will not be repeated here.

[0120] Here, the contrast loss function The goal is to maximize the similarity between positive samples and minimize the similarity between negative samples.

[0121] In the embodiment of the present invention, an optimization algorithm such as Adam or SGD is used to calculate the contrast loss function. Update model parameters:

[0122] in, are model parameters, including all weights and biases that need to be optimized. During the training process, the model is adjusted To minimize the contrast loss function . Represents the learning rate, which is used to control the step size of parameter updates. Denotes the contrast loss function For model parameters The gradient of , where “=” means assignment or update.

[0123] In some embodiments, in order to enhance the adaptability of the model when processing Chinese text, a large amount of Chinese corpus is introduced. Feature extraction is performed on the Chinese text samples, and the same model architecture is used to ensure that the image and text are represented in the same feature space.

[0124] Through the embodiment of the present invention, a one-stage contrast loss training is performed to enable the cross-modal image-text retrieval model to maximize the similarity between positive samples and minimize the similarity between negative samples.

[0125] According to a cross-modal image-text retrieval processing method provided by the present invention, after performing contrast loss training on the cross-modal image-text retrieval model based on the contrast loss function, the method further includes: Get query sample pairs; Determine a difficult positive sample pair and a difficult negative sample pair corresponding to the query sample pair, wherein the distance between the difficult positive sample pair and the query sample pair in the feature space is greater than a distance threshold, and the distance between the difficult negative sample pair and the query sample pair in the feature space is less than the distance threshold; Based on the query sample pairs, the difficult positive sample pairs and the difficult negative sample pairs, image triples and text triples are constructed; Determine query image features, positive sample image features, and negative sample image features of the image triplet; Determine query text features, positive sample text features, and negative sample text features of the text triples; Determine a triplet loss function based on the query image features, the positive sample image features, the negative sample image features, the query text features, the positive sample text features, and the negative sample text features; The cross-modal image-text retrieval model is trained with triplet loss based on the triplet loss function, where the triplet loss function is: ; in, represents the triplet loss function, Indicates selecting the maximum value. represents the distance function, represents the query image features, Represents the positive sample image features, Represents the negative sample image features, Indicates the preset first boundary value, Represents query text features, Represents the positive sample text features, Represents negative sample text features, Indicates the preset second boundary value.

[0126] After the first stage, samples that perform poorly in similarity calculation are selected as difficult samples. These samples are usually in the fuzzy area of ​​the model classification boundary, which makes it difficult for the model to accurately judge the similarity. Specifically, there are two types of difficult samples: one is difficult positive samples, which belong to the same category as the anchor samples, but have large differences in feature representation, resulting in lower-than-expected similarity calculation values; the other is difficult negative samples, which belong to different categories from the anchor samples, but are extremely close in feature space, resulting in higher similarity calculation values.

[0127] For each query sample pair , select a difficult positive sample pair and a hard negative pair , forming an image triplet or text triples .

[0128] Ensure that the selected difficult positive sample pairs are close to the query sample pairs in the feature space (less than the distance threshold), while the difficult negative sample pairs are far from the query sample pairs in the feature space (greater than the distance threshold).

[0129] For samples that constitute image triples or text triples, extract image features and text features:

[0130]

[0131]

[0132] in, represents the query image feature, represents the positive sample image feature, represents the negative sample image feature, represents the query text feature, represents the positive sample text feature, represents the negative sample text feature, represents an image sample of a query sample pair, An image sample representing a difficult positive pair, An image sample representing a difficult negative pair, A text sample representing a query sample pair, Text samples representing difficult positive pairs, Text samples representing difficult negative pairs, Represents a Transformer-based model, Represents a pre-trained language model.

[0133] In this embodiment of the present invention, the triple loss function Used to optimize the model, the goal is to make the distance between the anchor sample pair (query sample pair) and the difficult positive sample pair smaller than the distance between the anchor sample pair and the difficult negative sample pair, and the gap between the two must be larger than the preset boundary value. Triplet loss function Please refer to the above formula for details, and the present invention will not go into details here.

[0134] in, Represents a distance function (such as Euclidean distance):

[0135] in, represents the distance function, represents the first input, Represents the second input.

[0136] Using optimization algorithms such as gradient descent, according to the triple loss function Update model parameters:

[0137] in, are model parameters, including all weights and biases that need to be optimized. During the training process, the model is adjusted To minimize the triple loss function . Represents the learning rate, which is used to control the step size of parameter updates. Represents the triple loss function For model parameters The gradient of , where “=” means assignment or update.

[0138] Through the embodiment of the present invention, through the above two stages of training, the present invention can effectively improve the performance of cross-modal image-text retrieval. In the first stage, the model embedding space is optimized through contrast loss, so that similar images and texts are closer in the embedding space, and non-similar samples are farther away. In the second stage, the training of difficult samples is further strengthened through triple loss, so that the model is more robust when facing complex queries.

[0139] The present invention proposes a complete solution around cross-modal image and text retrieval, aiming to improve the accuracy and performance of retrieval. First, the present invention provides a process for generating an external knowledge base and realizing cross-modal image and text retrieval. The input document is segmented, and the text and image are encoded separately, and then the external knowledge base is generated through modal fusion. After the user's question is encoded, the similarity is calculated with the features in the knowledge base, and the most relevant search results are returned. Depending on the result type (text, image, or both), accurate answers are given through large language models or visual language models.

[0140] In terms of modal fusion methods, the present invention optimizes the fusion process of image and text information. Image features are extracted based on the Transformer model, and text features are extracted through a pre-trained language model. The weighted average and cross-modal attention mechanisms are used to fuse information from different modalities and generate a unified multimodal embedding representation, thereby better capturing the correlation information between images and texts and significantly improving the accuracy and performance of retrieval.

[0141] The present invention proposes a two-stage training method for cross-modal image and text retrieval. In the first stage, contrast loss is used to optimize the model embedding space by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. In particular, a large amount of Chinese corpus is introduced to enhance the adaptability of the model in processing Chinese text and visual information. In the second stage, triplet loss is used for training on difficult samples that performed poorly in the first stage, and negative samples are introduced to further optimize the embedding space, so that the distance between positive samples and query samples is closer, and the distance between negative samples and query samples is farther. Through this innovative training method, the present invention can better capture the semantic association between images and texts, significantly improve the effect of cross-modal image and text retrieval, and enhance the model's ability to process multilingual data.

[0142] The cross-modal image-text retrieval processing system provided by the present invention is described below. The cross-modal image-text retrieval processing system described below and the cross-modal image-text retrieval processing method described above can be referenced to each other.

[0143] refer to Figure 4 , Figure 4 It is a module schematic diagram of the cross-modal image and text retrieval processing system provided by the present invention.

[0144] The acquisition module 401 is used to acquire the query text input by the user during the image and text retrieval process.

[0145] The encoding module 402 is used to encode the query text through a text encoder to generate a query text feature vector.

[0146] The similarity module 403 is used to perform similarity matching based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base through a cross-modal image-text retrieval model, and return relevant results corresponding to a target number of multimodal embedding representations greater than a matching threshold, wherein the multimodal embedding representation is used to represent the joint features of the image and the text, and the relevant results include at least one of the following: image and text.

[0147] The retrieval result module 404 is used to input the relevant results and the query text into a preset multimodal large model when the relevant results include both images and texts, perform image question answering with text assistance, and obtain the retrieval results output by the multimodal large model.

[0148] Specifically, the above-mentioned cross-modal image and text retrieval processing system provided by the present invention can implement all the method steps implemented by the above-mentioned cross-modal image and text retrieval processing method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail here.

[0149] Figure 5 is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as Figure 5As shown, the electronic device may include: a processor 510 (processor), a communication interface 520 (Communications Interface), a memory 530 (memory) and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a cross-modal image-text retrieval processing method, the method comprising: obtaining a query text input by a user during the image-text retrieval process; encoding the query text through a text encoder to generate a query text feature vector; performing similarity matching based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base through a cross-modal image-text retrieval model, and returning a target number of related results corresponding to the multimodal embedding representation greater than the matching threshold, wherein the multimodal embedding representation is used to represent the joint features of the image and the text, and the related results include at least one of the following: image and text; when the related results include both the image and the text, the related results and the query text are input into a preset multimodal large model, and image question and answer with text assistance is performed to obtain the retrieval results output by the multimodal large model.

[0150] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0151] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the cross-modal image and text retrieval processing method provided by the above methods, the method including: obtaining a query text input by a user during the image and text retrieval process; encoding the query text through a text encoder to generate a query text feature vector; through a cross-modal image and text retrieval model, based on the query text feature vector and a multimodal embedding representation stored in an external knowledge base, similarity matching is performed, and a target number of related results corresponding to the multimodal embedding representation greater than a matching threshold are returned, wherein the multimodal embedding representation is used to represent the joint features of an image and text, and the related results include at least one of the following: image and text; when the related results include both an image and text, the related results and the query text are input into a preset multimodal large model, and image question and answer with text assistance is performed to obtain a retrieval result output by the multimodal large model.

[0152] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the cross-modal image-text retrieval processing method provided by the above-mentioned methods, the method comprising: obtaining a query text input by a user during an image-text retrieval process; encoding the query text through a text encoder to generate a query text feature vector; performing similarity matching based on the query text feature vector and a multimodal embedding representation stored in an external knowledge base through a cross-modal image-text retrieval model, and returning a target number of related results corresponding to multimodal embedding representations greater than a matching threshold, wherein the multimodal embedding representation is used to represent the joint features of an image and text, and the related results include at least one of the following: an image and text; when the related results include both an image and text, inputting the related results and the query text into a preset multimodal large model, performing image question and answering with text assistance, and obtaining a retrieval result output by the multimodal large model.

[0153] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0154] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross-modal image and text retrieval processing method, characterized in that: include: Obtain the query text entered by the user during the image and text retrieval process; Encoding the query text by a text encoder to generate a query text feature vector; Through a cross-modal image-text retrieval model, similarity matching is performed based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base, and relevant results corresponding to the multimodal embedding representations of a target number greater than a matching threshold are returned, wherein the multimodal embedding representation is used to represent the joint features of the image and the text, and the relevant results include at least one of the following: image and text; When the relevant results include both images and texts, the relevant results and the query text are input into a preset multimodal macro model to perform image question answering with text assistance to obtain the retrieval results output by the multimodal macro model.

2. The cross-modal image-text retrieval processing method according to claim 1, characterized in that: Before performing similarity matching based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base, the method further includes: Perform image-text segmentation on the input document to obtain text data and image data; Encoding the image data by an image encoder to obtain an input image feature vector; Encoding the text data by a text encoder to obtain an input text feature vector; Based on the cross-modal attention mechanism, the input text feature vector and the input image feature vector are modally fused to obtain a multimodal embedding representation; An external knowledge base is constructed based on the multimodal embedding representation.

3. The cross-modal image-text retrieval processing method according to claim 2, characterized in that: The image encoder is a Transformer-based model, and encoding the image data by the image encoder to obtain an input image feature vector includes: Dividing the image data into a plurality of image blocks of fixed size; Mapping each image block of the multiple image blocks to an embedding space of fixed length to obtain an embedding vector of each image block; Performing position encoding on each of the image blocks to obtain a relative position of each of the image blocks; According to the self-attention mechanism, based on the embedding vector of each image block and the relative position of each image block, an input image feature vector corresponding to the image data is determined.

4. The cross-modal image-text retrieval processing method according to claim 2, characterized in that: The method of performing modal fusion on the input text feature vector and the input image feature vector based on the cross-modal attention mechanism to obtain a multi-modal embedding representation includes: Determine a first attention weight between the i-th input image feature vector and the j-th input text feature vector; Determining a second attention weight between the j-th input text feature vector and the i-th input image feature vector; Determine a first fusion result of the input image feature vector on the input text feature vector based on the first attention weight; Determine a second fusion result of the input text feature vector on the input image feature vector based on the second attention weight; The first fusion result and the second fusion result are combined to obtain a multimodal embedding representation.

5. The cross-modal image-text retrieval processing method according to claim 1, characterized in that: Before performing similarity matching based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base through the cross-modal image-text retrieval model and returning relevant results corresponding to the multimodal embedding representations with a target number greater than a matching threshold, the method further includes: Obtaining an image-text sample pair training set, wherein the image-text sample pair training set includes positive sample pairs and negative sample pairs; Determining a first similarity between the image feature and the text feature of the positive sample pair and a second similarity between the image feature and the text feature of the negative sample pair; Determining a contrast loss function based on the first similarity and the second similarity; The cross-modal image-text retrieval model is trained with contrast loss based on the contrast loss function, wherein the contrast loss function is: ; in, represents the contrast loss function, represents the total number of sample pairs of the image-text sample pair training set, represents the index of the positive sample pair, Indicates The first similarity of positive sample pairs, represents the index of the negative sample pair, Indicates selecting the maximum value. Indicates The second similarity of the negative sample pairs, Indicates a preset first similarity threshold.

6. The cross-modal image-text retrieval processing method according to claim 5, characterized in that: After performing contrast loss training on the cross-modal image-text retrieval model based on the contrast loss function, the method further includes: Get query sample pairs; Determine a difficult positive sample pair and a difficult negative sample pair corresponding to the query sample pair, wherein a distance between the difficult positive sample pair and the query sample pair in the feature space is greater than a distance threshold, and a distance between the difficult negative sample pair and the query sample pair in the feature space is less than the distance threshold; Constructing image triples and text triples based on the query sample pair, the difficult positive sample pair, and the difficult negative sample pair; Determine query image features, positive sample image features, and negative sample image features of the image triplet; Determine query text features, positive sample text features, and negative sample text features of the text triples; Determining a triplet loss function based on the query image feature, the positive sample image feature, the negative sample image feature, the query text feature, the positive sample text feature, and the negative sample text feature; The cross-modal image-text retrieval model is trained with triple loss based on the triple loss function, wherein the triple loss function is: ; in, represents the triple loss function, Indicates selecting the maximum value. represents the distance function, represents the query image feature, represents the positive sample image feature, represents the negative sample image feature, Indicates the preset first boundary value, represents the query text feature, represents the positive sample text feature, represents the negative sample text feature, Indicates the preset second boundary value.

7. A cross-modal image and text retrieval processing system, characterized in that: include: The acquisition module is used to obtain the query text entered by the user during the image and text retrieval process; An encoding module, used for encoding the query text through a text encoder to generate a query text feature vector; A similarity module is used to perform similarity matching based on the query text feature vector and the multimodal embedding representation stored in the external knowledge base through a cross-modal image-text retrieval model, and return relevant results corresponding to the multimodal embedding representations of a target number greater than a matching threshold, wherein the multimodal embedding representation is used to represent the joint features of the image and the text, and the relevant results include at least one of the following: image and text; The retrieval result module is used to input the relevant results and the query text into a preset multimodal large model when the relevant results include both images and texts, perform image question answering with text assistance, and obtain the retrieval results output by the multimodal large model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the cross-modal image and text retrieval processing method as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-modal image and text retrieval processing method as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the cross-modal image and text retrieval processing method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Text modal and image modal crossing type data retrieval method

    CN105205096A

  • Multi-modal medical report retrieval method and device, electronic equipment and storage medium

    CN116645554A

  • Cross-modal molecular information retrieval method, system, equipment and medium

    CN118069900A

  • Visual question and answer method based on knowledge retrieval enhancement

    CN119513245A

Cited By

  • Security human-vehicle multi-mode retrieval method and device, electronic equipment and storage medium

    CN120234439A

  • Security person and vehicle multi-modal retrieval method and device, electronic equipment and storage medium

    CN120234439B

  • Geological data retrieval method and device, electronic device and storage medium

    CN120296213A

  • Fault diagnosis and suggestion method based on interpretable artificial intelligence and large model

    CN120338140A

  • Cross-modal data query method

    CN120632128A