Lightweight and efficient multi-modal document retrieval enhancement method

By training a hybrid retrieval architecture that combines sparse encoders and multi-vector encoders with sparse indexes and multi-vector representations, the problems of recognition error and cumbersome operation in OCR technology are solved, achieving lightweight and efficient multimodal document retrieval enhancement, and improving retrieval accuracy and response efficiency.

CN121009219AActive Publication Date: 2025-11-25ZHEJIANG UNIV +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511508169.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-25
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Among existing multimodal document retrieval methods, OCR technology suffers from large recognition errors and severe loss of visual and spatial structural information, resulting in inaccurate retrieval results. Furthermore, it is cumbersome and time-consuming to operate, making it difficult to achieve efficient and accurate multimodal document retrieval enhancement.

Method used

We train a sparse encoder, a multi-vector encoder, and a multi-modal document reorderer using a multimodal large language model. Combining sparse retrieval with a multi-vector fine-grained interaction strategy, we introduce a time-cost-aware loading mechanism to build a lightweight and efficient multimodal document retrieval system. Through a hybrid retrieval architecture of sparse index and multi-vector representation, we optimize memory and computation efficiency.

Benefits of technology

It significantly improves the accuracy and response efficiency of multimodal document retrieval, solves the memory bottleneck and response latency problems in traditional methods, and achieves efficient and accurate multimodal document retrieval enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009219A_ABST
    Figure CN121009219A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight and efficient multi-modal document retrieval enhancement method, which comprises the following steps of: training a sparse encoder, a multi-vector encoder and a multi-modal document rearrangement device by utilizing a multi-modal large language model, and combining sparse retrieval and a multi-vector fine granularity delay interaction strategy; by designing a mixed retrieval architecture combining sparse index and multi-vector representation, the problems of memory bottleneck and response delay faced by a traditional multi-modal retrieval method in a mass data scene are effectively relieved; an introduced time cost perception loading strategy and a clustering storage mechanism collaboratively optimize the delay interaction score calculation efficiency; and finally, sparse and delayed interaction scores are fused, and a multi-modal language model resorter is combined, so that the document matching precision is remarkably improved, and the response efficiency of a multi-modal retrieval enhancement system and the user experience are comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multimodal information retrieval technology, specifically relating to a lightweight and efficient method for enhancing multimodal document retrieval. Background Technology

[0002] With the rapid development of big data and artificial intelligence technologies, multimodal documents (such as PDF, Word, and PPT documents, which contain text, images, charts, and other modal information) are widely used in industries such as law, healthcare, finance, and insurance, becoming an indispensable information resource in data lakes and data warehouses. Accurately locating relevant documents from massive multimodal documents using natural language queries is a crucial foundation for achieving Retrieval-Augmented Generation (RAG) tasks. Multimodal document retrieval enhancement methods (such as the technology in the Chinese patent application with publication number CN119988588A) typically first accurately retrieve multimodal documents related to the user's query, and then input these document contents along with the user's query into a large language model to generate an answer, thereby effectively improving the accuracy and credibility of the generated results.

[0003] Currently, multimodal document retrieval enhancement technologies typically rely on optical character recognition (OCR) technology and traditional text retrieval techniques. Specifically, these methods first convert visual content such as images and charts in a document into text using OCR technology (e.g., in the literature [Wang B, Xu C, Zhao X et al. Mineru: An open-source solution for accurate document content extraction [J]. arXiv preprint: 2409.18839, 2024]), and then use traditional text retrieval methods (such as BM25) or dense vector retrieval methods for subsequent retrieval processing.

[0004] In the process of realizing this invention, we have discovered at least the following problems in the prior art: On the one hand, OCR technology itself has a certain degree of recognition error, especially when processing complex charts and images. On the other hand, visual and spatial structural information is inevitably lost during OCR conversion, making it difficult to fully capture the complete semantics of a document during retrieval. These factors significantly reduce the accuracy of retrieval results and make it difficult to effectively support subsequent augmentation generation tasks. In addition, the OCR process is cumbersome and time-consuming, usually requiring the use of specialized recognition models for different modalities, such as dedicated table recognition models or formula recognition models, which further increases the difficulty and cost of implementation.

[0005] Therefore, there is an urgent need to propose a lightweight multimodal document retrieval enhancement method and system that can ensure retrieval accuracy while possessing a simple structure and high processing efficiency, so as to effectively address the pressing need for efficient and accurate retrieval enhancement of large-scale multimodal documents in practical applications. Summary of the Invention

[0006] In view of the above, the present invention provides a lightweight and efficient multimodal document retrieval enhancement method. By utilizing a multimodal large language model to train a sparse encoder, a multi-vector encoder and a multimodal document reorderer, and combining sparse retrieval with a multi-vector fine-grained post-interaction strategy, and introducing a time cost-aware loading mechanism, the method significantly reduces memory and computational overhead while improving the accuracy and response efficiency of multimodal document retrieval.

[0007] A lightweight and efficient multimodal document retrieval enhancement method includes the following steps: (1) Obtain a multimodal document library, in which the documents contain one or more modal information including text, images, charts, and tables; (2) Based on the open-source dataset and the multimodal large language model, train a sparse encoder, a multi-vector encoder and a multimodal document reordering machine; (3) Encode the documents in the multimodal document library using the trained sparse encoder and multivector encoder, and extract the sparse vector representation and dense multivector representation of the documents respectively; (4) Construct an inverted index structure based on sparse vector representation and store it on disk. Use a disk storage algorithm based on balanced clustering to store dense multi-vector representation; (5) Load the inverted index structure and use sparse retrieval to quickly filter out candidate documents and calculate the sparse score; for candidate documents, use a time cost-aware loading strategy to dynamically load the dense multi-vector representation of the candidate documents from the disk and calculate the delayed interaction score; further filter the candidate documents based on the score after fusing the sparse score and the delayed interaction score. (6) The candidate documents retained after filtering are combined with the user's natural language query through a template and then input into the trained multimodal document reorderer for reordering to obtain the most relevant documents; (7) Input the natural language query and the most relevant document into the multimodal large language model to automatically generate questions and answers and return the final results to the user.

[0008] Furthermore, the documents in the multimodal document library are in any format, including PDF, Word, and PPT, and each document consists of structured metadata (such as title, chapter level, author information), natural language paragraphs, illustrations, charts, tables, and formulas.

[0009] Furthermore, the sparse encoder in step (2) is used to learn a sparse representation with high compression ratio suitable for large-scale inverted index construction; the multi-vector encoder is used to capture the fine-grained semantic structure inside the document to support subsequent token-level delayed interactive retrieval; the multimodal document reorderer is used to perform high-precision semantic matching and sorting of candidate documents; the LoRA (low-rank adaptive) lightweight parameter fine-tuning update strategy is adopted in the training process of the above three.

[0010] Further, the specific implementation of step (3) is as follows: First, the documents in the multimodal document library are uniformly converted into PDF format, and the documents are image-processed at the page level to generate image sequences; then, the image sequences are sequentially input into the trained sparse encoder and multi-vector encoder for encoding. The sparse encoder generates a bag-of-words sparse representation by projecting the token-level hidden states obtained from the image encoding, performing ReLU (corrected linear unit) function and max pooling, which is used to construct an inverted index structure and support preliminary document retrieval; the multi-vector encoder retains all the token hidden states output by the image encoding and uniformly compresses them into a 128-dimensional vector set through linear mapping to form a fine-grained dense multi-vector representation for subsequent token-level fine-grained delayed interactive retrieval.

[0011] Furthermore, in step (4), for the sparse vector representation of each document, only the weights of the non-zero dimensions are retained, and the document IDs corresponding to the non-zero dimensions are mapped to the weights to form a sparse inverted index, thereby constructing an inverted index structure from the keyword dimension to the document set; at the same time, using the sparse vector representation as a feature, all documents are divided into several semantically similar and equally sized clusters through a balanced clustering algorithm, and then the dense multi-vector representation of the document is written to the disk according to the clustering results in a continuous physical location, ensuring that the dense multi-vector representation of the document in the same cluster is stored in the same disk, so as to reduce the randomness of disk access during reading.

[0012] Further, the specific implementation of step (5) is as follows: First, the user's natural language query is encoded using a sparse encoder and a multi-vector encoder to obtain sparse query vectors and dense query multi-vectors; then, the inverted index structure is loaded, and the sparse score between the sparse query vector and the sparse vector representation of each document is calculated based on the vector overlap dimension. The top M documents with the highest sparse scores are selected as candidate documents; then, a time cost-aware loading strategy is adopted to load the dense multi-vector representation of the candidate documents. When there are many candidate documents in a cluster, the dense multi-vector representation of the entire cluster is loaded; if there are only a few candidate documents in a cluster, only the dense multi-vector representation of the candidate documents is loaded; the delayed interaction score between the dense query multi-vector and the dense multi-vector representation of each candidate document is calculated based on the token-level delayed interaction matching similarity algorithm. Finally, the fusion score of each candidate document is calculated according to the following formula, and the candidate documents are further filtered according to the fusion score. Where: score is the fusion score of the candidate documents, score spa The sparse score for candidate documents, score mul Let Z be the delayed interaction score of the candidate document, Z() be the standardization function, α be the weight coefficient, and M be a natural number greater than 1.

[0013] Further, in step (6), each candidate document retained after filtering is concatenated with the user's natural language query in the order of "query-candidate document" to form an input sequence. These input sequences are then input into the multimodal document reorderer one by one. The reorderer projects the hidden layer state corresponding to the last token of the input sequence onto the vocabulary space of the multimodal large language model and obtains the log probability of "Yes" as the relevance score. Then, the candidate documents are reordered according to the relevance score, and the candidate document with the highest relevance score is selected as the most relevant document.

[0014] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described multimodal document retrieval enhancement method.

[0015] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described multimodal document retrieval enhancement method.

[0016] This invention effectively alleviates the memory bottleneck and response latency issues faced by traditional multimodal retrieval methods in massive data scenarios by designing a hybrid retrieval architecture that combines sparse indexes and multi-vector representations. The introduced time-cost-aware loading strategy and clustering storage mechanism synergistically optimize the computational efficiency of delayed interaction. Finally, by fusing sparse and delayed interaction scores and combining them with a multimodal language model reorderer, the document matching accuracy is significantly improved, and the response efficiency and user experience of the multimodal retrieval enhancement system are comprehensively improved. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the lightweight and efficient multimodal document retrieval enhancement method of the present invention. Detailed Implementation

[0018] To describe the present invention in more detail, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] The multimodal document retrieval enhancement scheme proposed in this invention is driven by a multimodal large language model. Based on the trained sparse vector encoder and multi-vector encoder, it transforms documents into structurally clear and semantically rich sparse representations and fine-grained multi-vector representations. A fusion algorithm is then used for efficient representation matching and retrieval score integration. Simultaneously, this invention introduces a disk-based contiguous storage mechanism for document clustering and a time-cost-aware loading strategy, reducing memory usage and access latency while ensuring computational efficiency during the delayed interaction score calculation stage. Finally, a reorderer driven by the multimodal large language model performs high-precision reordering of candidate documents, significantly improving the final retrieval relevance. Overall, it achieves the design goals of "storage saving, latency optimization, and accuracy improvement," and is suitable for efficient retrieval enhancement tasks in various scenarios, including large-scale RAG and multimodal intelligent systems.

[0020] Example like Figure 1 As shown, this embodiment provides a lightweight and efficient multimodal document retrieval enhancement method, the specific implementation process of which is as follows: Step S11: Obtain a multimodal document library, where the document format can be PDF, Word, PPT, etc., and the document contains one or more modal information such as text, images, charts, and tables.

[0021] Specifically, this invention can be widely applied to various practical scenarios such as government information disclosure, scientific research literature management, corporate archives, medical record searches, legal and regulatory access, and patent examination assistance. Without loss of generality, the multimodal document library in this embodiment is represented as a dataset containing heterogeneous documents, sourced from publicly available internet resources, corporate knowledge bases, or scientific research databases. Each document consists of structured metadata (such as title, chapter hierarchy, and author information), natural language paragraphs, illustrations, charts, tables, and formulas. For example, in a scientific research literature scenario, a document may contain embedded experimental visualizations, LaTeX formulas, and structured result tables.

[0022] Step S12: Train a sparse encoder, a multi-vector encoder, and a multimodal reorderer based on an open-source dataset and a multimodal large language model to achieve efficient representation and refined sorting of documents.

[0023] Specifically, in this embodiment, each sample in the open-source dataset contains a query and a corresponding document page for model training. A unified encoding and ranking framework is constructed using a multimodal large language model, with independent optimization objectives designed for three types of sub-modules: the sparse encoder aims to learn sparse representations with high compression rates suitable for large-scale inverted index construction; the multi-vector encoder is used to capture fine-grained semantic structures within documents to support subsequent token-level delayed interactive retrieval; and the multimodal reorderer is used for high-precision semantic matching and ranking of candidate documents. To improve training efficiency and enhance deployment flexibility, each sub-module adopts a low-rank adaptive fine-tuning lightweight parameter update strategy, significantly reducing training resource overhead while ensuring the model's generalization ability in multiple scenarios.

[0024] The specific training process is as follows: The sparse encoder is optimized using contrastive learning loss. After inputting the document image into the multimodal large language model, the extracted hidden states are mapped to the vocabulary space, and sparse features are extracted through ReLU activation and max pooling operations, ultimately generating a bag-of-words representation for sparse retrieval. The multi-vector encoder retains all token-level hidden states generated by image encoding and further compresses them into a set of vectors of a uniform dimension (e.g., 128-dimensional) through linear mapping to form a dense vector representation that supports lazy loading and interactive computation. Its training objective is also to minimize the contrastive learning loss. The reorderer is trained based on an instruction fine-tuning method, constructing a multimodal input sequence containing the query and the document image, and optimizing the model's ranking ability by maximizing the log probability of the "Yes" token. The above three sub-modules share the underlying multimodal large language base model and efficiently complete their respective functions in GPU memory through a modular switching mechanism, thereby constructing a collaboratively optimized multimodal retrieval system.

[0025] Step S13: First, convert the multimodal documents to PDF format and convert each page of PDF to an image. Then, use a sparse encoder and a multi-vector encoder to encode the documents in a multimodal manner, and obtain sparse vector representations and multi-vector representations that can be used for sparse retrieval and delayed interaction score calculation, respectively.

[0026] Specifically, in this embodiment, the original multimodal documents (including PDF, Word, PPT and other formats) are first uniformly converted into PDF format to ensure the consistency of input format and the fidelity of visual structure; then, each PDF document is processed into images at the page level to generate a high-resolution image page sequence, which serves as the subsequent visual input source.

[0027] Then, encoding is performed. The system inputs each page's image in parallel into the sparse encoder and multi-vector encoder, which were trained in step S12, to extract two types of complementary representation information. The sparse encoder projects, ReLU activation, and max-pooling the token-level hidden states obtained from the image encoding to generate a bag-of-words sparse vector representation. This sparse vector is used to construct an inverted index structure and support initial document retrieval. Simultaneously, the multi-vector encoder retains all token hidden states output by the image encoder and compresses them uniformly into a 128-dimensional vector set through a linear mapping module, forming a fine-grained dense semantic multi-vector representation for subsequent token-level fine-grained delayed interactive retrieval. This process supports generating independent vector encodings for each page of the document, and the vectors from all pages can be aggregated to form the global representation structure of the entire document.

[0028] It should be noted that the entire preprocessing and encoding process can be executed efficiently in parallel on GPUs (graphics processing units) or dedicated acceleration devices, supporting batch document vectorization processing, and is decoupled from the subsequent index building and retrieval modules, allowing for independent deployment and updates, significantly improving the system's maintainability and engineering availability.

[0029] Step S14: Construct an efficient and scalable inverted index structure based on sparse vector representation, use sparse vectors to cluster documents, and use the clustering results to optimize the disk storage of multi-vector representation to improve loading efficiency.

[0030] Specifically, the extracted sparse vector representation is expressed as: Where i represents the i-th document page, |V| is the vocabulary size of the multimodal large language model, and w ji This represents the importance weight of a document on the j-th token. Only sparse vector entries with non-zero dimensions are retained, and an inverted index structure is constructed from the keyword dimension to the document set: This inverted index structure supports fast retrieval of candidate document sets during the online query phase using the non-zero dimension of the sparse query vector, and calculates a sparse score based on the vector overlap dimension. ).

[0031] Furthermore, to reduce the disk I / O cost in the subsequent multi-vector delayed interaction score calculation stage, this embodiment further adopts a document clustering strategy based on sparse vector semantic distribution to optimize the sequential disk storage of multi-vector representations. Specifically, using document sparse vectors as input features, an improved balanced clustering algorithm divides documents into several semantically similar and equally sized clusters. The algorithm ensures that the number of documents within each cluster meets the balance constraint, avoiding the loading performance instability caused by extremely large or small clusters. Subsequently, the system writes the multi-vector representations of the documents sequentially to the disk based on the clustering results, ensuring that the multi-vector representations of documents within the same cluster are physically and contiguously stored in the same disk block. This allows for efficient batch sequential loading with less random I / O cost after subsequent sparse retrieval to select candidate documents.

[0032] Step S15: In the retrieval stage, firstly, sparse retrieval is used to quickly filter out candidate documents. Then, for the filtered candidate documents, according to the time cost-aware reading algorithm, the full cluster loading or the multi-vector loading of the specified document is dynamically selected to perform refined delayed interaction calculation, and the sparse score and the multi-vector delayed interaction score are fused. The top-K relevant documents are selected according to the ranking.

[0033] Specifically, firstly, based on the constructed inverted index structure, and by matching the non-zero dimensions of the sparse vector with the keywords in the inverted index, an efficient initial document screening operation is performed on the user query. Specifically, the query input is encoded into a sparse vector using a sparse encoder. The candidate document set is quickly retrieved using the non-zero dimensions of the sparse query vector, and a sparse score is calculated based on the vector overlap dimension. ).

[0034] After selecting the top-M highest-scoring candidate documents in the sparse retrieval phase, the user query input is first encoded into a set of multi-vector representations using a multi-vector encoder to capture fine-grained semantic information in the query. Then, a dynamic loading strategy is determined for the multi-vector representation of each of the M candidate documents. This strategy comprehensively considers the number of candidate documents, disk read / write throughput, and cluster size. Specifically, based on the distribution of the vector storage clusters to which the candidate documents belong on disk, the system comprehensively considers the number of documents within the cluster, the proportion of the target document within the cluster, and the sequential and random read efficiency of the storage device. Two loading strategies are adopted for each candidate document's storage cluster: one is "whole cluster loading," where when the cluster contains many target documents, the vector data of the entire cluster is loaded sequentially into memory at once, reducing the number of random disk accesses and improving loading efficiency; the other is "specified loading," where when the cluster contains only a few target documents, the multi-vector representation corresponding to the target document is directly located and read from the disk, avoiding unnecessary redundant data loading. This time-cost-aware dynamic loading strategy ensures that the overall efficiency of document multi-vector loading is maximized while maintaining retrieval accuracy.

[0035] After vector loading is complete, the system performs token-based delayed interaction matching similarity score calculation for each query q and candidate document, and calculates multi-vector delayed interaction scores. and the obtained sparse score The scores from multiple vector delayed interactions are normalized and then fused to obtain the final comprehensive document score, fully leveraging the complementary advantages of sparse and dense data. The final document ranking is then determined based on the fused scores. The fusion formula for this process is as follows: Where: α is the weighting coefficient, and Z() is the standardization function. For sparse retrieval scores, For multi-vector delayed interaction scores, the fusion formula is used to balance the impact of the two types of scores on the final result.

[0036] Step S16: Concatenate the top-K relevant documents with the user query according to the sorting template, input them into the multimodal large language model reorderer, use the log probability of the "Yes" word output by the last token of the sequence as the final relevance score, and further reorder the documents.

[0037] Specifically, this embodiment employs a multimodal large language model based on instruction fine-tuning as a rearranger and supports multimodal input, including query text and document image content. In practice, the system first concatenates each "query-candidate document" pair into a complete input sequence according to a predefined Prompt template. This sequence explicitly guides the model to determine the relevance between the document and the query. For example, the input template could be "Given the user query: <query>, isthe following document relevant? <document>The "Answer Yes or No." part consists of an image embedding or encoded representation of the document and a query part consisting of the user's natural language input.

[0038] The concatenated sequence is then fed into a multimodal reorderer, which, based on its judgment ability learned during the training phase, predicts the log probability of the "Yes" token from the input sequence as the final relevance score between the candidate document and the current query. The system performs the same process on all candidate documents and performs a final reordering according to the score to select the document results that best match the user's intent semantically.

[0039] Through the above operations, step S16 achieves high-precision reordering based on prior sparsity and delayed interaction, effectively improving the document retrieval system's adaptability to modal heterogeneity and semantically complex scenarios. This step lays a precise document input foundation for the final question-answering generation task.

[0040] Step S17: Input the user query and the most relevant documents into the multimodal large language model to achieve various downstream tasks such as question answering and summarizing, and return the final results to the user to improve the user's interactive experience.

[0041] The above description of the embodiments is provided to enable those skilled in the art to understand and apply the present invention. Those skilled in the art can readily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made to the present invention by those skilled in the art based on the disclosure thereof should be within the scope of protection of the present invention.< / document> < / query>

Claims

1. A lightweight and efficient multimodal document retrieval enhancement method, characterized in that, Includes the following steps: (1) Obtain a multimodal document library, in which the documents contain one or more modal information including text, images, charts, and tables; (2) Based on the open-source dataset and the multimodal large language model, train a sparse encoder, a multi-vector encoder and a multimodal document reordering machine; (3) Encode the documents in the multimodal document library using the trained sparse encoder and multivector encoder, and extract the sparse vector representation and dense multivector representation of the documents respectively; (4) Construct an inverted index structure based on sparse vector representation and store it on disk. Use a disk storage algorithm based on balanced clustering to store dense multi-vector representation; (5) Load the inverted index structure and use sparse retrieval to quickly filter out candidate documents and calculate the sparse score; for candidate documents, use a time cost-aware loading strategy to dynamically load the dense multi-vector representation of the candidate documents from the disk and calculate the delayed interaction score. Candidate documents are further filtered based on the combined score of sparse score and delayed interaction score; (6) The candidate documents retained after filtering are combined with the user's natural language query through a template and then input into the trained multimodal document reorderer for reordering to obtain the most relevant documents; (7) Input the natural language query and the most relevant document into the multimodal large language model to automatically generate questions and answers and return the final results to the user.

2. The multimodal document retrieval enhancement method according to claim 1, characterized in that: The documents in the multimodal document library can be in any format, including PDF, Word, and PPT. Each document consists of structured metadata, natural language paragraphs, illustrations, charts, tables, and formulas.

3. The multimodal document retrieval enhancement method according to claim 1, characterized in that: The sparse encoder in step (2) is used to learn a sparse representation with high compression ratio that is suitable for large-scale inverted index construction; The multi-vector encoder is used to capture the fine-grained semantic structure within the document to support subsequent token-level delayed interactive retrieval; the multimodal document reorderer is used to perform high-precision semantic matching and ranking of candidate documents; the LoRA lightweight parameter fine-tuning update strategy is adopted during the training of the above three components.

4. The multimodal document retrieval enhancement method according to claim 1, characterized in that: The specific implementation of step (3) is as follows: First, the documents in the multimodal document library are uniformly converted into PDF format, and the documents are image-processed at the page level to generate image sequences; then, the image sequences are sequentially input into the trained sparse encoder and multi-vector encoder for encoding. The sparse encoder generates a bag-of-words sparse representation by projecting, using the ReLU function and max pooling to the token-level hidden states obtained from the image encoding, which is used to construct an inverted index structure and support preliminary document retrieval; the multi-vector encoder retains all the token hidden states output by the image encoding and compresses them uniformly into a 128-dimensional vector set through linear mapping to form a fine-grained dense multi-vector representation for subsequent token-level fine-grained delayed interactive retrieval.

5. The multimodal document retrieval enhancement method according to claim 1, characterized in that: In step (4), for the sparse vector representation of each document, only the weights of the non-zero dimensions are retained, and the document IDs corresponding to the non-zero dimensions are mapped to the weights to form a sparse inverted index, thereby constructing an inverted index structure from the keyword dimension to the document set; at the same time, using the sparse vector representation as a feature, all documents are divided into several semantically similar and equally sized clusters through a balanced clustering algorithm, and then the dense multi-vector representation of the document is written to the disk according to the clustering results in a continuous physical location, ensuring that the dense multi-vector representation of the document in the same cluster is stored on the same disk.

6. The multimodal document retrieval enhancement method according to claim 1, characterized in that, The specific implementation of step (5) is as follows: First, the user's natural language query is encoded using a sparse encoder and a multi-vector encoder to obtain sparse query vectors and dense query multi-vectors; then, the inverted index structure is loaded, and the sparse score between the sparse query vector and the sparse vector representation of each document is calculated based on the vector overlap dimension. The top M documents with the highest sparse scores are selected as candidate documents; then, a time cost-aware loading strategy is used to load the dense multi-vector representation of the candidate documents. If a cluster contains a large number of candidate documents, the dense multi-vector representation of the entire cluster is loaded; if a cluster contains only a small number of candidate documents, only the dense multi-vector representation of the candidate documents is loaded; the delayed interaction score between the dense query multi-vector and the dense multi-vector representation of each candidate document is calculated based on a token-level delayed interaction matching similarity algorithm. Finally, the fusion score of each candidate document is calculated according to the following formula, and the candidate documents are further filtered based on the fusion score. Where: score is the fusion score of the candidate documents, score spa The sparse score for candidate documents, score mul Let Z be the delayed interaction score of the candidate document, Z() be the standardization function, α be the weight coefficient, and M be a natural number greater than 1.

7. The multimodal document retrieval enhancement method according to claim 1, characterized in that: In step (6), each candidate document retained after filtering is concatenated with the user's natural language query in the order of "query-candidate document" to form an input sequence. These input sequences are then input into the multimodal document reorderer one by one. The reorderer projects the hidden layer state corresponding to the last token of the input sequence onto the vocabulary space of the multimodal large language model and obtains the log probability of "Yes" as the relevance score. Then, the candidate documents are reordered according to the relevance score, and the candidate document with the highest relevance score is selected as the most relevant document.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: The processor is used to execute the computer program to implement the multimodal document retrieval enhancement method as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the multimodal document retrieval enhancement method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal document retrieval enhancement generation method based on large model

    CN119988588A

  • RAG-based vertical domain knowledge multi-round question and answer method

    CN118964556A

  • Multi-modal retrieval enhancement generation technology-based vertical industry document intelligent analysis and question answering method

    CN119474482A

  • Retrieval enhancement generation method and system based on multi-source knowledge fusion and performance verification method

    CN119621935A

  • Video retrieval generation method and device based on sparse representation and reordering

    CN120316308A