Multi-modal retrieval method and system based on fine-grained interaction

By employing a fine-grained interactive multimodal retrieval method, documents are converted into images and multi-dimensional embeddings are generated. Cross-modal fusion is then performed using an alignment network, which solves the problem that traditional retrieval systems struggle to handle multimedia documents and achieves more efficient and accurate information retrieval.

CN120910283AActive Publication Date: 2025-11-07MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511020155.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Traditional information retrieval systems struggle to effectively integrate and process complex documents containing multiple media formats, resulting in insufficient accuracy and relevance of retrieval results.

Method used

We employ a multimodal retrieval method based on fine-grained interaction. By converting documents into images and generating page embeddings and text summaries, we extract high-dimensional features using a multimodal model, perform multi-dimensional embedding of images and text, and map them into the same latent space through an alignment network. Combining coarse-grained and fine-grained retrieval strategies, we achieve cross-modal fusion analysis.

Benefits of technology

It significantly improves the accuracy and relevance of multimodal information retrieval, enabling a more comprehensive understanding and representation of information in complex documents, reducing information omissions and misjudgments, and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910283A_ABST
    Figure CN120910283A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal retrieval method and system based on fine-grained interaction, and the method comprises the steps: converting each page of a document into an image, adding a text abstract, maintaining the context relationship between the image and the text, enabling the text and the image content to be integrated under the same frame, and improving the retrieval efficiency. And the integrity and consistency of the information in the conversion process are ensured. The page image is analyzed, multi-dimensional embedding of each area is generated, and it is ensured that information of each area can be accurately extracted and processed. By performing fine-grained embedding generation on each region, semantic information contained in different parts in the image can be accurately understood, and the semantic information is combined with text information to ensure efficient matching of the information. Through multi-dimensional embedding and fine-grained interactive alignment, the features of the image and the text are mapped to the same hidden space, so that the correlation with finer granularity is captured. According to the method, the deep relation between the image and the text can be deeply mined, and the information retrieval and generation precision and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of multi-modal document retrieval, and specifically relates to a multi-modal retrieval method and system based on fine-grained interaction. BACKGROUND

[0002] With the rapid development of information technology, the generation and dissemination of multimedia data have become increasingly widespread. Modern documents are no longer limited to pure text, but increasingly contain various media forms such as tables, images, etc. The emergence of such multi-modal documents poses new challenges to information retrieval and processing technology. Traditional information retrieval systems are mainly based on text data, and are difficult to comprehensively understand and process complex documents containing multiple media forms, resulting in insufficient accuracy and relevance of the retrieval results. With the diversification of document content, especially the introduction of non-text information such as images, charts, audio, etc., existing technologies are difficult to effectively integrate and process such heterogeneous information, and there is an urgent need for new methods to improve the understanding ability of information and retrieval effect. SUMMARY

[0003] The purpose of the present application is to explore and implement efficient multi-modal information retrieval technology, improve the accuracy and comprehensiveness of information retrieval, and multi-modal information retrieval can more comprehensively understand and represent the information in complex documents, by fusing text, table and image data, to provide more relevant and comprehensive retrieval results, and reduce information omission and misjudgment.

[0004] In order to achieve the above purpose, the present application proposes a multi-modal retrieval method based on fine-grained interaction, comprising:

[0005] Each page of the document to be queried is converted into an image, and a page embedding model is used to generate a page embedding and a corresponding text summary for each image; the image is parsed into multiple logical regions, and a multi-modal model is used to generate multi-dimensional text and image embeddings for each logical region;

[0006] High-dimensional features are extracted from the image using a multi-modal model to generate multi-dimensional embeddings of the image; word embedding or sentence embedding technology is used to convert user query text into multi-dimensional text embeddings; the multi-dimensional embeddings of the image and the text are decomposed into multiple image tokens and text tokens; the correlation score between each pair of image token and text token embeddings is calculated to form an interaction matrix, each element in the interaction matrix corresponding to the similarity of a token pair; the scores in the matrix are aggregated to generate a final global relevance score; an alignment network is used to map image tokens and text tokens to the same hidden space; by training the alignment network, the mapping function is optimized so that the features of the image and the text are as close as possible in the hidden space;

[0007] The similarity between the text embedding of the user query text and the page embedding is calculated, the semantic filtering of the page text summary is combined, the first N most possible page images containing the search result are screened as candidate pages, so as to narrow the range of subsequent search;

[0008] Based on the multi-dimensional embedding in the logical region and the alignment network of the candidate page, the similarity between the user query text and the multi-dimensional embedding of each logical region is calculated, the key region in the candidate page is further screened in combination with the type characteristics of the logical region, and the most relevant local content set to the user query text is returned through cross-modal fusion analysis.

[0009] As an improvement of the above method, the page embedding model is ColQwen2, Qwen2.5-VL-32B or VisRAG-Ret.

[0010] As an improvement of the above method, the image is parsed into multiple logical regions using the MinerU tool.

[0011] As an improvement of the above method, the multi-modal model is CLIP or ViT-BERT.

[0012] As an improvement of the above method, the word embedding or sentence embedding technology is BERT or RoBERTa.

[0013] As an improvement of the above method, the cosine similarity or dot product is used to calculate the correlation score between each pair of image token and text token embedding.

[0014] As an improvement of the above method, the alignment network is SBANet or MagNet.

[0015] The application also provides a multi-modal retrieval system based on fine-grained interaction, which is realized based on the above method, and the system comprises:

[0016] A document semantic fine-grained representation module is used to convert each page of a document to be queried into an image, generate a page embedding and a corresponding text summary of each image by using a page embedding model, parse the image into multiple logical regions, and generate multi-dimensional text and image embedding of each logical region by using a multi-modal model.

[0017] The cross-modal interaction module based on image-text alignment is used for extracting high-dimensional features from images by using a multi-modal model to generate multi-dimensional embedding of the images, converting user query text into multi-dimensional embedding of the text by using word embedding or sentence embedding technology, decomposing the multi-dimensional embedding of the images and the text into multiple image tokens and text tokens, calculating the correlation score between each pair of image token and text token embedding to form an interaction matrix, mapping the image tokens and the text tokens into the same hidden space by using an alignment network, and optimizing the mapping function by training the alignment network so that the features of the images and the text are as close as possible in the hidden space.

[0018] The coarse-grained retrieval module is used for calculating the similarity of the text embedding of the user query text and the page embedding, combining the semantic filtering of the page text summary, and screening out the top N page images most likely to contain the retrieval results as candidate pages.

[0019] The fine-grained retrieval module is used for calculating the similarity of the user query text and the multi-dimensional embedding of each logical region based on the multi-dimensional embedding of the logical regions of the candidate pages and the alignment network, further screening out the key regions in the candidate pages by combining the type characteristics of the logical regions, and returning the local content set most relevant to the user query text by cross-modal fusion analysis.

[0020] Compared with the prior art, the application has the following advantages:

[0021] 1. The application designs a set of document semantic fine-grained representation method, which includes two stages of retrieval process, the first stage is coarse-grained retrieval, and the relevant page images are retrieved according to the user input query; the second stage is fine-grained retrieval, and the text and image content in the page images retrieved in the first stage are further retrieved.

[0022] 2. The application designs a set of image-text alignment cross-modal interaction method, which uses multi-dimensional embedding technology to extract high-dimensional features of images and texts respectively, and generates rich multi-modal embedding vectors; then, the image and text features are aligned by the alignment network to ensure the consistency and integrity of the information. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 Fig. 1 shows a flowchart of a multi-modal retrieval method based on fine-grained interaction;

[0024] Figure 2 Fig. 2 shows a structural diagram of a document semantic fine-grained representation technology;

[0025] Figure 3The structure diagram of cross-modal interaction based on text-image alignment is shown; wherein, PositionalEmbedding: position embedding; Textual token: text token; Visual token: image token; WeightedSum: weighted sum. DETAILED DESCRIPTION

[0026] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings.

[0027] Multimodal information retrieval technology can comprehensively understand and represent information content by utilizing data of various media forms such as text, image, video, and audio, thereby significantly improving the accuracy and relevance of information retrieval. Compared with traditional text retrieval methods, multimodal retrieval can better capture the relationship and semantics between different media forms and reduce the understanding bias caused by single or missing information. In addition, multimodal information retrieval can better meet the user's demand for diversified information, provide more rich and intuitive search results, and significantly improve the user experience. For documents containing complex elements such as pictures and charts, traditional retrieval methods often cannot effectively extract and match relevant information, while multimodal information retrieval can comprehensively process these elements to ensure higher retrieval accuracy.

[0028] At the same time, the development of this technology will also promote the technological progress in related fields such as natural language processing, computer vision, and audio processing, and promote the wide application of these technologies. Through deeper cross-domain collaboration, multimodal technology can enhance the interaction between various fields and promote the further development of multimodal learning and understanding. For example, in the fields of medical treatment, education, entertainment, and security, multimodal information retrieval can help users more accurately obtain the required information and promote the popularization and landing of intelligent applications.

[0029] Multimodal information retrieval technology significantly improves the accuracy and comprehensiveness of information retrieval, enhances the quality and diversity of information generation, improves the level of automation and intelligence, and improves the user experience. It not only provides a new idea for traditional information retrieval systems, but also provides strong technical support for various complex multimodal data processing. These advantages make multimodal retrieval technology have wide practical application value, especially in processing large document libraries, news recommendation, cross-platform information retrieval, etc., which shows great potential.

[0030] To address the limitations of the prior art, the present application proposes a multi-modal retrieval method based on fine-grained cross-modal interaction. This method converts each page of a document into a single image and attaches a detailed text summary, maintaining the contextual relationship between images and text. In this way, text and image content can be integrated under the same framework, ensuring the integrity and consistency of information during conversion and processing. Further, the page image is parsed to generate multi-dimensional embeddings for each region, ensuring that information in each sub-region can be accurately extracted and processed. By generating fine-grained embeddings for each region, the system can accurately understand the semantic information contained in different parts of the image and combine it with text information, ensuring efficient matching of information.

[0031] Through multi-dimensional embedding and fine-grained interaction alignment technology, the features of images and text are mapped to the same hidden space, further realizing fine-grained cross-modal interaction, thereby capturing more fine-grained relevance. This method can deeply explore the deep relationship between images and text, improving the accuracy and efficiency of information retrieval and generation. In addition, the introduction of fine-grained interaction enables the system to capture more subtle differences and relationships when associating image and text information, significantly enhancing the accuracy and relevance of the retrieval results.

[0032] Through the combination of these technologies, the present application not only effectively processes multi-modal documents, but also provides more efficient and accurate solutions in the field of multi-modal information retrieval. This method fully integrates image and text information, improving the system's cross-modal understanding ability, and has wide application prospects and market value.

[0033] The present application proposes a multi-modal retrieval method based on fine-grained interaction, which first processes the document to be retrieved based on document semantic fine-grained, then performs cross-modal processing based on image-text alignment, and finally performs semantic retrieval on the document. This method includes:

[0034] Step 1) Fine-grained representation technology based on document semantics;

[0035] Step 2) Cross-modal interaction technology based on image-text alignment;

[0036] Step 3) Two-stage retrieval process from coarse to fine.

[0037] In the above technical solution, step 1) specifically includes:

[0038] Step 101) Convert each page of the document into a complete page image, which preserves the overall layout and visual information of the page. Through this conversion, the structured information such as typography, font, color, and graphic-text relationship in the page is completely captured, allowing subsequent processing to conduct fine-grained analysis at the image level, ensuring that each part of the information in the document is reasonably expressed and understood.

[0039] Step 102) For each page image, generate page embeddings and corresponding text summaries using now mainstream page embedding models (such as ColQwen2, VisRAG-Ret) so that the language model can better understand the content of the image. The generation of text summaries can rely on the visual model Qwen2.5-VL-32B, which can effectively extract key information from the image and provide more concise and explicit text data for subsequent retrieval tasks.

[0040] Step 103) Parse the page image into multiple logical regions such as title area, text area, chart area, list area, etc. using tools such as MinerU. The content characteristics and structural information of each region, including text, images, and tables, are extracted and analyzed separately, providing a basis for subsequent cross-modal alignment and embedding generation. This process not only helps to extract the diversity of document content, but also enables each part of the information in the document to be processed and analyzed accurately.

[0041] Step 104) Apply multi-modal models (such as CLIP, ViT-BERT, etc.) to each parsed fine-grained element (such as text, image, table) to generate multi-dimensional text and image embeddings for each region. Through these models, image regions and text information can be mapped into a common embedding space, and these embeddings can capture complex semantic relationships between regions and provide effective support for subsequent cross-modal retrieval.

[0042] Step 105) Use efficient vector storage techniques (such as FAISS, etc.) to quickly store and retrieve the embeddings of each page and fine-grained element, reducing computational and storage overhead; dynamically adjust the division strategy of the region and the generation method of the embedding according to the needs of the actual application scenario, and flexibly cope with different document types and complexity. Efficient storage and retrieval techniques such as FAISS can significantly improve the running efficiency of the system on large-scale data sets, and through dynamic adjustment strategies to adapt to various complex situations, further optimize the retrieval process.

[0043] In the above technical solution, step 2) specifically includes:

[0044] Step 201) Encode image and text information with multi-dimensional embedding instead of simplifying them into a single one-dimensional embedding. Through multi-dimensional embedding, image and text information can be more accurately expressed, ensuring fine-grained understanding of information and effective cross-modal docking. This processing method can better capture the diversity and complexity of information, avoiding the loss of information in traditional methods. Specific steps include: using multi-modal models (such as CLIP, ViT-BERT, etc.) to extract high-dimensional features from images to generate multi-dimensional embeddings of images. Each feature dimension in the image corresponds to its different semantic levels, so that the embedding representation can cover the diverse information in the image.

[0045] Step 202) Convert query text into multi-dimensional vector representation using word embedding or sentence embedding technology (such as BERT, RoBERTa, etc. pre-training model). These pre-training models can capture complex semantic information and contextual relationships of text through deep learning methods, laying the foundation for effective fusion of images and text. The embedding of the text can effectively map the lexical, grammatical, and semantic information in the text, generating a more rich representation of the text.

[0046] Step 203) Decompose the multi-dimensional embeddings of images and texts into multiple tokens, each representing a local feature or word. This step further subdivides the embedding representation of images and texts into independent parts through fine-grained tokenization, facilitating the processing of local features one by one. In this way, the system can better capture the subtle connections between images and texts and improve the accuracy of cross-modal retrieval through precise tokenization processing.

[0047] Step 204) Achieve fine-grained semantic alignment between images and texts by constructing an interaction matrix: first, calculate the correlation score between each pair of image-text token embeddings using cosine similarity or dot product to form a complete interaction matrix; then, aggregate the scores in the matrix through maximum value, average value or attention weighting to generate the final global relevance score. Each element in the matrix corresponds to the similarity of a specific token pair, while the aggregation operation integrates these local matching signals into a quantitative indicator of overall semantic fit, preserving both microscopic semantic associations and macroscopic matching evaluations.

[0048] Step 205) By introducing an alignment network (such as SBANet, MagNet, etc.), the visual features (image tokens) and the language features (text tokens) are mapped into the same latent space, so that they can be compared and fused in the same space. The alignment network optimizes the mapping function by learning to make the features of images and texts as close as possible in the latent space, so that they can be more seamlessly fused. This mapping process improves the interaction and understanding depth between images and texts.

[0049] Step 206) By training the alignment network, the mapping function is optimized so that the features of images and texts are as close as possible in the latent space. The alignment network maps the visual features and the language features into the same latent space, so that they can be more effectively compared and fused. This optimization process can improve the accuracy of cross-modal retrieval, so that different modal data can work better together.

[0050] In the above technical solution, step 3) specifically includes:

[0051] A two-stage retrieval strategy is adopted to realize accurate information positioning from the whole to the local: the first stage of coarse-grained page retrieval takes the user query as input, calculates the similarity (such as cosine similarity) between the query text embedding and the page embedding generated in step 102), and combines the semantic filtering of the page text summary (such as Qwen3-8B-Reranker model, etc.) to quickly filter out the top-N most likely to contain answers Page image, narrowing the search range; the second stage of fine-grained area retrieval directly targets these candidate pages, based on the logical areas (title, body, chart, etc.) divided in step 103) and the multi-dimensional embedding of the areas generated based on the cross-modal interaction technology in step 2), the similarity between the query text and the area embedding is calculated, and the area type features (such as preferentially matching the "chart area" to the query containing data visualization) are combined to further filter out the key areas within the page, and finally through cross-modal fusion analysis (such as associating chart visual features with body text description) to return the most relevant local content set. This progressive retrieval strategy from page to area not only ensures efficiency through coarse-grained retrieval, but also improves accuracy through fine-grained analysis, effectively balancing the search range and the accuracy of information positioning.

[0052] Through these steps, the method of the present application can effectively improve the efficiency and accuracy of multi-modal retrieval, especially in handling complex semantic relationships between images and texts, it shows strong processing capacity. This retrieval method based on fine-grained cross-modal interaction has wide application prospects in the field of information retrieval.

[0053] Embodiment 1

[0054] As Figure 1As shown, the embodiment utilizes a fusion retrieval technology of multi-dimensional embedding and fine-grained interaction, combines high-dimensional features based on images with word embedding or sentence embedding based on text, and balances and integrates the contributions of different modal features, so as to more comprehensively model the relationship between queries and documents and enhance the ability of the retrieval model in multi-modal information understanding and relevance determination. Specifically, by converting each page of the document into a single image and attaching a detailed text summary, the context relationship between the image and the text is maintained. Further, the page image is parsed to generate multi-dimensional embedding of each region. At the same time, an alignment network (such as a multi-layer perception MLP or a convolutional neural network CNN) is introduced to map visual features and language features to the same hidden space, realize fine-grained cross-modal interaction, and capture finer-grained relevance. Through the combination of these technologies, the multi-modal document semantic retrieval system can more effectively process complex multi-modal documents and provide more accurate and comprehensive retrieval and generation results.

[0055] As shown, Figure 2 The document semantic fine-grained representation method provided by the embodiment includes:

[0056] Step 1) Preliminary processing of page images

[0057] In this stage, first, each page of the document is converted into a complete page image, which preserves the overall layout and visual information of the page. The image processing can combine traditional text information and visual information to provide a unified format and content for subsequent processing steps. Then, for each page image, a page embedding and a corresponding text summary are generated to enable the language model to better understand the content of the image. The text summary, as a supplement to the image information, can enhance the system's ability to capture important information in the image, making the retrieval task more accurate and comprehensive.

[0058] Step 2) Region division of page images

[0059] The page image is parsed into multiple logical regions, such as title area, text area, chart area, list area, etc.; the content type of each region is identified by tools such as MinerU, such as pure text, image, table, etc. Region division not only helps to structure the document content, but also finely extracts the information features of each region to provide more detailed data support for subsequent retrieval tasks. The accuracy of each region identification directly affects the quality of the retrieval results, so this step is an important part of the entire system.

[0060] Step 3) Efficient vector storage based on FAISS

[0061] By using efficient vector storage technologies such as FAISS, the embedding of each page and fine-grained element in the document can be quickly stored and retrieved, effectively reducing the computational and storage overhead. At the same time, combined with the actual application scene demand, dynamically adjusting the region division strategy and embedding generation method can flexibly adapt to different document types and complexity. These methods not only significantly improve the running efficiency of the system on large-scale data sets, but also optimize the retrieval process through dynamic adjustment to better cope with complex situations.

[0062] As shown in Figure 3 The text-alignment cross-modal interaction method provided by the embodiment includes:

[0063] Step 1) Multi-dimensional embedding

[0064] Instead of simplifying image and text information into a single one-dimensional embedding, multi-dimensional embedding is used to encode them. The specific steps include: using a multi-modal model (such as CLIP, ViT-BERT, etc.) to extract high-dimensional features from images and generate multi-dimensional embeddings of images; using word embedding or sentence embedding technology (such as BERT, RoBERTa, etc. pre-training model) to convert text into multi-dimensional vector representation. Through these high-dimensional representations, not only can the information of images and texts be more accurately captured, but also more semantic and detailed information can be retained, avoiding the information loss in traditional methods. Combining the multi-dimensional embeddings of images and texts, a comprehensive multi-modal embedding representation is formed, which can more comprehensively represent the document content and improve the accuracy and efficiency of retrieval.

[0065] Step 2) Fine-grained interaction

[0066] When calculating the correlation score, the interaction between each pair of token embeddings is considered, including cross-modal interaction. The specific steps include: decomposing the multi-dimensional embeddings of images and texts into multiple tokens, each token representing a local feature or word; constructing an interaction matrix to calculate the similarity or correlation score between each pair of token embeddings; by aggregating the scores in the interaction matrix, the final correlation score is obtained. Through fine-grained interaction, more fine-grained relevance can be captured, improving the accuracy and relevance of retrieval, and thus providing more accurate matching results when processing complex multi-modal documents. This step effectively improves the system's understanding of document content, so that each detail is fully considered.

[0067] Step 3) Alignment network

[0068] By introducing an alignment network, visual features and language features are mapped into the same latent space, so that they can be compared and fused in the same space. The specific steps include: using the alignment network to map the multidimensional embeddings of images and texts into the same latent space respectively; in the latent space, the features of images and texts can be directly compared and fused; by training the alignment network, the mapping function is optimized so that the features of images and texts are as close as possible in the latent space. The alignment network maps visual features and language features into the same latent space, so that they can be more effectively compared and fused, further improving the performance of the system. This process can significantly improve the integration capability of cross-modal information, ensure seamless docking between information, and further improve the intelligence and accuracy of the retrieval system.

[0069] As shown in Figure 2 The coarse-to-fine retrieval process provided by the embodiment includes:

[0070] The first stage is coarse-grained retrieval, which retrieves relevant page images according to the user input query; the second stage is fine-grained retrieval, which further retrieves the text and image content in the page images based on the page images retrieved in the first stage. Finally, accurate local information is returned. This progressive retrieval from the whole to the local not only guarantees efficiency, but also realizes accurate positioning of detailed information.

[0071] Through the combination of these technologies, the application provides an efficient and accurate multi-modal retrieval method, which can provide more accurate matching results in a complex document environment. This method can not only handle standard documents, but also cope with the comprehensive retrieval of various types of data such as images, charts, and texts, and has a wide application prospect.

[0072] Embodiment 2

[0073] The application also provides a multi-modal retrieval system based on fine-grained interaction, which is realized based on the above method, and the system includes:

[0074] A document semantic fine-grained representation module is used to convert each page of the to-be-queried document into an image, generate a page embedding and a corresponding text summary of each image using a page embedding model, and parse the image into multiple logical regions, and generate multidimensional text and image embeddings of each logical region using a multi-modal model.

[0075] The cross-modal interaction module based on image-text alignment is used for extracting high-dimensional features from images by using a multi-modal model to generate multi-dimensional embeddings of the images; converting user query text into multi-dimensional embeddings of the text by using word embedding or sentence embedding technology; decomposing the multi-dimensional embeddings of the images and the text into multiple image tokens and text tokens; calculating a correlation score between each pair of image token and text token embeddings to form an interaction matrix; mapping the image tokens and the text tokens into the same hidden space by using an alignment network; and optimizing a mapping function by training the alignment network so that the features of the images and the text are as close as possible in the hidden space.

[0076] The coarse-grained retrieval module is used for calculating the similarity between the text embedding of the user query text and the page embedding, combining the semantic filtering of the page text summary, and screening the top N page images most likely to contain the retrieval results as candidate pages, so as to narrow the scope of subsequent retrieval.

[0077] The fine-grained retrieval module is used for calculating the similarity between the user query text and the multi-dimensional embedding of each logical region based on the multi-dimensional embedding of the logical regions of the candidate pages and the alignment network, combining the type characteristics of the logical regions to further screen the key regions in the candidate pages, and returning the most relevant local content set to the user query text by cross-modal fusion analysis.

[0078] The application can also provide a computer device, which includes at least one processor, memory, at least one network interface and user interface. The various components in the device are coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between the components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus.

[0079] The user interface can include a display, a keyboard or a clicking device. For example, a mouse, a trackball, a touchpad or a touch screen, etc.

[0080] It can be appreciated that the memory in the embodiments disclosed in the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (Read-Only Memory, ROM), a programmable read-only memory (Programmable ROM, PROM), an erasable programmable read-only memory (Erasable PROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM) or a flash memory. The volatile memory can be a random access memory (Random Access Memory, RAM) used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (Static RAM, SRAM), dynamic random access memory (Dynamic RAM, DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (Synchlink DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory described herein is intended to include but not limited to these and any other suitable types of memory.

[0081] In some embodiments, the memory stores elements, executable modules or data structures, or a subset thereof, or an extended set thereof: an operating system and an application program.

[0082] Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program includes various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The program for implementing the method of the embodiments of the present disclosure can be included in the application program.

[0083] In the above-described embodiments, the processor can be configured to perform the steps of the above-described method by invoking the program or instructions stored in the memory, in particular, the program or instructions stored in the application program.

[0084] performing the steps of the above-described method.

[0085] The method can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip having a signal processing capability. In implementation, the steps of the method can be completed by an integrated logic circuit of hardware in the processor or by an instruction in the form of software. The processor can be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods disclosed above can be implemented or executed by the processor. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed above can be directly embodied as a hardware code executed by the processor or a combination of hardware and software modules in the processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage media is located in the storage memory, and the processor reads information in the storage memory and combines the hardware to complete the steps of the method.

[0086] It can be understood that the embodiments described in the present application can be implemented in hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing functions described in the present application or a combination thereof.

[0087] For software implementation, the present application can be implemented by executing functional modules (such as processes, functions, etc.) described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0088] The application further provides a nonvolatile storage medium for storing the computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and all of them should be covered in the scope of the claims of the present application.

Claims

1. A multi-modal retrieval method based on fine-grained interaction, comprising: converting each page of a document to be queried into an image, generating a page embedding and a corresponding text summary of each image using a page embedding model; parsing the image into multiple logical regions, generating multi-dimensional text and image embeddings of each logical region using a multi-modal model; extracting high-dimensional features from the image using the multi-modal model to generate multi-dimensional embeddings of the image; converting user query text into multi-dimensional embeddings of text using word embedding or sentence embedding technology; decomposing multi-dimensional embeddings of images and text into multiple image tokens and text tokens; calculating the correlation score between each pair of image token and text token embeddings to form an interaction matrix, each element of the interaction matrix corresponding to the similarity of a token pair; mapping image tokens and text tokens into the same hidden space using an alignment network; optimizing the mapping function by training the alignment network so that the features of images and text are as close as possible in the hidden space; calculating the similarity between the text embedding of the user query text and the page embedding, combining the semantic filtering of the page text summary to filter out the top N most likely to contain search results page images as candidate pages, thereby narrowing the scope of subsequent search; based on the multi-dimensional embeddings of the logical regions and the alignment network in the candidate pages, calculating the similarity between the user query text and the multi-dimensional embeddings of each logical region, combining the type characteristics of the logical regions to further filter out the key regions within the candidate pages, and returning the most relevant local content set to the user query text through cross-modal fusion analysis.

2. The method of claim 1, wherein, The page embedding model is ColQwen2, Qwen2.5-VL-32B or VisRAG-Ret. 3.The method of claim 1, wherein, The image is parsed into multiple logical regions using the MinerU tool. 4.The method of claim 1, wherein, The multi-modal model is CLIP or ViT-BERT.

5. The method of claim 1, wherein, The word embedding or sentence embedding technology is BERT or RoBERTa. 6.The method for multi-modal retrieval based on fine-grained interaction according to claim 1, characterized in that, The correlation score between each pair of image token and text token embeddings is calculated using cosine similarity or dot product. 7.The method for multi-modal retrieval based on fine-grained interaction according to claim 1, characterized in that, The alignment network is SBANet or MagNet.

8. A multi-modal retrieval system based on fine-grained interaction, implemented based on the method of any one of claims 1-7, characterized in that, The system comprises: a document semantic fine-grained representation module for converting each page of a document to be queried into an image, generating a page embedding and a corresponding text summary of each image using a page embedding model, and parsing the image into multiple logical regions, generating multi-dimensional text and image embeddings of each logical region using a multi-modal model; a cross-modal interaction module based on image-text alignment for extracting high-dimensional features from the image using a multi-modal model to generate multi-dimensional embeddings of the image, converting user query text into multi-dimensional embeddings of text using word embedding or sentence embedding technology, decomposing multi-dimensional embeddings of images and text into multiple image tokens and text tokens, calculating the correlation score between each pair of image token and text token embeddings to form an interaction matrix, and mapping image tokens and text tokens into the same hidden space using an alignment network; A coarse-grained retrieval module is configured to calculate the similarity between the text embedding of the user query text and the page embedding, combine the semantic filtering of the page text summary, and screen out the top N most likely to contain the retrieval result page images as candidate pages, thereby narrowing the scope of subsequent retrieval. A fine-grained retrieval module is configured to calculate the similarity between the user query text and the multi-dimensional embedding of each logical region based on the multi-dimensional embedding in the logical region and the alignment network of the candidate page, combine the type characteristics of the logical region to further screen out the key region in the candidate page, and return the most relevant local content set to the user query text through cross-modal fusion analysis.

Citation Information

Patent Citations

  • Image-text retrieval method and system based on semantic information reasoning and cross-modal interaction

    CN118133839A

  • Capacitance detection report document retrieval method based on text image alignment

    CN118606498A

  • Cross-modal image text retrieval method based on deep learning

    CN119311911A

  • Remote sensing cross-modal retrieval method based on deep ternary fusion sensing network

    CN119336968A

  • File retrieval and management method fusing AI large model and graph data

    CN119884038A