Multi-modal document retrieval enhancement generation method based on large model

Through the enhanced generation method of multimodal document retrieval based on large models, the problems of difficulty in data parsing, low long text processing efficiency and insufficient accuracy of search results in multimodal document processing are solved, and efficient and accurate multimodal document retrieval and generation are achieved, improving the reliability of user experience and information.

CN119988588APending Publication Date: 2025-05-13杭州长望智创科技有限公司
View PDF 0 Cites 64 Cited by

Patent Information

Application Number
CN202510155596.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-12
Filing Date
2025-02-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art faces problems such as difficulty in multimodal data parsing, low long text processing efficiency and insufficient accuracy of search results when processing multimodal documents.

Method used

The multimodal document retrieval enhancement generation method based on large models is adopted, and user queries and multimodal documents are processed through natural language processing and embedded models, query vectors and document vectors are generated, and hierarchical index structure is constructed to achieve efficient retrieval and generation.

Benefits of technology

It improves the comprehensiveness and accuracy of multimodal data analysis and processing, improves the accuracy and relevance of search results, improves the user experience, and ensures the real-time and reliability of information through dynamic update mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988588A_ABST
    Figure CN119988588A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multi-modal data processing, and particularly relates to a multi-modal document retrieval enhancement generation method based on a large model, which comprises the following steps: receiving query content input by a user for a multi-modal document; processing the query content by adopting an embedded model, generating a query vector representing user query semantic information, and storing the query vector in a vector database; analyzing the multi-modal document to obtain long text information, segmenting the long text information into data blocks by adopting a recursive partitioning strategy, and numbering and marking the data blocks; carrying out vectorization processing on the data blocks by adopting an embedded model to generate document vectors, storing the document vectors into a vector database, and constructing a hierarchical index structure; retrieving in a vector database based on the query vector, and returning a retrieval result; and processing a retrieval result by utilizing a large language model to generate response content conforming to the query intention of the user. According to the method, the multi-modal document can be effectively analyzed and processed, and the accuracy and comprehensiveness of analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal data processing, and in particular relates to a multimodal document retrieval enhancement generation method based on a large model. Background Art

[0002] With the rapid development of information technology, various document formats (such as PDF, plain text, code, Markdown, etc.) have been widely used in various fields. The existing RAG (Retrieval-Augmented Generation) technology faces the following challenges when processing these multimodal documents: Difficulty in parsing multimodal data: Traditional methods are difficult to process multiple types of data such as text, pictures, and tables at the same time, which affects the overall parsing effect; Low efficiency in processing long texts: The processing of long documents often leads to low retrieval and generation efficiency, affecting user experience; Insufficient accuracy of retrieval results: Existing single retrieval methods (such as vector retrieval or keyword retrieval) are difficult to meet the needs of complex queries, resulting in poor accuracy and relevance of retrieval results. Summary of the invention

[0003] Purpose of the invention: The purpose of the present invention is to address the deficiencies of the prior art and provide a large-model-based multimodal document retrieval enhancement generation method to effectively parse and process multimodal documents (including text, images, tables, etc.) and improve the accuracy and comprehensiveness of the analysis.

[0004] Technical solution: The multimodal document retrieval enhancement generation method based on a large model of the present invention comprises the following steps: S1. receiving a query content input by a user for a multimodal document, parsing the query content input by the user using natural language processing technology, and extracting semantic information; S2. Processing the query content using an embedded model to generate a query vector representing the user's query semantic information, and storing the query vector in a vector database; S3, parsing the multimodal document to obtain long text information, using a recursive block strategy to divide the long text information into data blocks that meet a predetermined size, and numbering the data blocks; S4, using an embedded model to vectorize the data block to generate a document vector and store it in a vector database to build a hierarchical index structure; S4, searching the vector database based on the query vector, and returning a search result; S5. Process the search results using a large language model to generate response content that meets the user's query intent; S6. Display the original metadata of the search results, including: document title, author, creation time, document source and version information, and document classification labels; dynamically update the vector database and keyword index according to changes in the data environment, including: monitoring the update status of the data source, detecting new or modified documents, updating the vector representation and index in real time, and maintaining data consistency.

[0005] To further improve the above technical solution, in step S1, the method of receiving the user query content includes text input method and voice input method; the text input method supports natural language text input, provides query association and completion functions, and performs input verification in real time; the voice input method uses voice recognition technology to convert voice, reduce noise and enhance signals, and transcribe voice content in real time; the query content is processed by natural language technology, including: intent recognition and entity extraction, identification of query keywords and semantic relationships, and construction of a structured representation of the query.

[0006] Furthermore, the multimodal document includes one or more formats of PDF, HTML, TXT, and Markdown; parsing the multimodal document includes: using image recognition technology to parse non-text elements in the document and converting them into structured data; using OCR technology to identify and extract text in the document, using natural language processing technology to perform grammatical analysis and entity recognition on the extracted text content, and organizing the parsed content into structured data according to a predetermined format.

[0007] Furthermore, the recursive chunking strategy includes: using predefined delimiters to recursively divide long text information into data blocks; when the first segmentation fails to meet the predetermined conditions, reprocessing the generated data blocks using different delimiters or segmentation criteria until data blocks that meet the required size or structural characteristics are obtained; assigning a unique identifier to each data block and recording its location information in the original document; establishing associations between data blocks, including contextual relationships and hierarchical relationships.

[0008] Furthermore, the vectorization processing includes: constructing a training data set, including: taking sentence pairs within the same paragraph as positive sample pairs, randomly selecting sentence pairs from different documents as negative sample pairs, and adding specific prefix tags for different types of data; training an embedding model, including: defining a contrast loss function based on cosine distance, optimizing model parameters using small batch stochastic gradient descent, and evaluating model performance using a validation set; generating a vector representation, including: word embedding the input text, processing the word embedding through a multi-layer Transformer encoder, and obtaining a contextual vector representation of the text; optimizing the vector representation, including: performing vector dimensionality reduction processing, normalizing the vector, and performing vector quality evaluation.

[0009] Furthermore, the vector database adopts a hierarchical index structure, including: the first-level index, including: building an index table containing the basic information of all documents, storing the summary information and keywords of the documents, and establishing the association relationship between documents; the second-level index, including: building a fine-grained index of the specific content of the document, recording the vector representation at the paragraph and sentence level, and storing the semantic relationship between data blocks.

[0010] Furthermore, the retrieval step includes: using multiple similarity calculation methods to calculate the similarity scores between the query vector and all vectors in the vector database; sorting the retrieval results based on the similarity scores; deduplicating the retrieval results to remove duplicate or highly similar content; reordering the retrieval results using a Transformer-based cross encoder; and combining the results of multiple similarity calculation methods to obtain the final ranking.

[0011] Furthermore, a variety of similarity calculation methods include: cosine similarity calculation: similarity is calculated by vector dot product, and vector norm is normalized; a similarity score in the range of [-1, 1] is obtained; Euclidean distance calculation: the sum of squares of differences in each dimension of the vector is calculated, the result is squared, and the distance is converted into a similarity score; Jaccard similarity calculation: the vector is converted into a set representation, the ratio of the intersection and the union of the sets is calculated, and a similarity score in the range of [0, 1] is obtained.

[0012] Furthermore, the large language model processing steps include: operating path planning for user queries, including: analyzing query type and complexity, determining query decomposition strategy, and selecting a combination of processing methods; decomposing complex queries into multiple sub-queries, including: identifying multiple semantic units in the query, establishing logical relationships between sub-queries, and executing sub-query processing in parallel; processing the context in blocks, including: dividing the retrieval results into blocks according to semantic units, analyzing semantic associations between blocks, and ensuring context coherence; generating multiple candidate answers, including: generating answers based on different context blocks, evaluating the relevance and accuracy of the answers, and selecting the best answer combination; integrating the final response, including: organizing the logical order of multiple answers, ensuring coherence between answers, and generating complete response content.

[0013] Beneficial effects: The present invention provides a retrieval enhancement generation method for accurate parsing and segmentation of multimodal documents based on a large model, which has significant advantages and improvements compared with the prior art, which are specifically reflected in the following aspects: 1. Comprehensiveness and accuracy of multimodal data processing The parsing module of the present invention can process documents in various formats (such as PDF, HTML, TXT, etc.) and convert non-text elements (such as pictures, tables) therein into structured data. This multimodal data processing method, combined with advanced image recognition and OCR technology, can fully and accurately extract information from documents. Traditional methods can usually only process a single data type, while the present invention can process text, images and tables at the same time, improving the integrity and accuracy of data parsing.

[0014] 2. Accurate and efficient recursive text segmentation strategy The present invention adopts a recursive block segmentation strategy to segment long texts. Specifically, the long texts are recursively divided into smaller data blocks using predefined delimiters (such as periods, commas, semicolons, etc.). If the first segmentation fails to meet the predetermined conditions, different delimiters or segmentation criteria will be used for reprocessing until data blocks that meet the requirements are obtained. This method ensures the accuracy and semantic integrity of text segmentation, and avoids the semantic fragmentation problem caused by improper segmentation in traditional methods.

[0015] 3. Efficient vectorization and hierarchical index structure The present invention vectorizes the data blocks by using an embedding model based on weakly supervised pre-training and contrastive learning, making the vector representation more accurate. In the vector database, a hierarchical index structure is used to first build an index containing all document summary information, and then build an index covering each part of the document in detail. This hierarchical structure greatly improves the retrieval efficiency and accuracy, and solves the problems of slow retrieval speed and inaccurate results caused by imperfect indexes in traditional methods.

[0016] 4. Intelligent retrieval and prompt word generation The present invention searches based on the user query vector and returns relevant documents, prompt words and query content. By calculating the similarity between the query vector and the vector in the database, and sorting, deduplicating and merging the results, a new text paragraph array is generated. The search results are reordered through a cross encoder to ensure that the results finally presented to the user are the most relevant and valuable. Compared with traditional keyword searches, when displaying the search results, the present invention not only provides the information summarized by the large model, but also displays the original metadata of the relevant documents, such as title, author, publication date, etc. This approach enhances the transparency and credibility of the information, and users can check the information generated by the model with the original data source to ensure the accuracy of the information.

[0017] 5. Large language models generate high-quality responses The present invention uses a large language model to further process the search results and generate high-quality response content. The specific steps include decomposing complex queries into multiple sub-queries for parallel execution, collecting and fusing them into coherent sentences, and generating the final reply. The large language model can analyze and optimize according to the context, extract more accurate and relevant answers, and overcome the problem of low quality of generated results in traditional methods.

[0018] 6. Accuracy and consistency of response synthesis In the response synthesis process, the present invention summarizes the retrieved context, screens and highlights key information, and lays the foundation for generating more accurate answers. Based on different context blocks, multiple targeted answers are generated, and these answers are integrated or summarized to form a final reply. This method ensures that the final generated response has high accuracy and coherence, and improves the user experience.

[0019] 7. Dynamic update mechanism The system of the present invention can dynamically update the vector database and keyword index according to changes in the data environment, ensuring that the system can always provide the latest and most relevant information. In traditional methods, untimely data updates may lead to outdated or inaccurate retrieval results, but the dynamic update mechanism of the present invention effectively solves this problem and improves the real-time performance and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of a retrieval enhancement generation method for accurate parsing and segmentation of multimodal documents based on a large model of the present invention; Figure 2 It is a schematic diagram of the multimodal document accurate classification and parsing method of the present invention; Figure 3 It is a schematic diagram of the method for recursive hierarchical segmentation of long text information of the present invention. DETAILED DESCRIPTION

[0021] The technical solution of the present invention is described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the embodiments.

[0022] Example 1: Figure 1 The multimodal document retrieval enhancement generation method based on a large model includes the following steps: S1. receiving a query content input by a user for a multimodal document, parsing the query content input by the user using natural language processing technology, and extracting semantic information; S2. Processing the query content using an embedded model to generate a query vector representing the user's query semantic information, and storing the query vector in a vector database; S3, parsing the multimodal document to obtain long text information, using a recursive block strategy to divide the long text information into data blocks that meet a predetermined size, and numbering the data blocks; S4, using an embedded model to vectorize the data block to generate a document vector and store it in a vector database to build a hierarchical index structure; S4, searching the vector database based on the query vector, and returning a search result; S5. Process the search results using a large language model to generate response content that meets the user's query intent; S6. Display the original metadata of the search results, including: document title, author, creation time, document source and version information, and document classification labels; dynamically update the vector database and keyword index according to changes in the data environment, including: monitoring the update status of the data source, detecting new or modified documents, updating the vector representation and index in real time, and maintaining data consistency.

[0023] Specifically, in step S1, the method of receiving user query content includes text input method and voice input method; the text input method supports natural language text input, provides query association and completion functions, and performs input verification in real time; the voice input method uses voice recognition technology to convert voice, reduce noise and enhance signals, and transcribe voice content in real time; the query content is processed by natural language technology, including: intent recognition and entity extraction, identification of query keywords and semantic relationships, and construction of a structured representation of the query.

[0024] Multimodal documents include one or more formats such as PDF, HTML, TXT, and Markdown; parsing multimodal documents includes: using image recognition technology to parse non-text elements in the document and converting them into structured data; using OCR technology to identify and extract text in the document, using natural language processing technology to perform syntax analysis and entity recognition on the extracted text content, and organizing the parsed content into structured data according to a predetermined format.

[0025] The recursive chunking strategy includes: using predefined delimiters to recursively divide long text information into data blocks; when the first segmentation fails to meet the predetermined conditions, reprocessing the generated data blocks using different delimiters or segmentation criteria until data blocks that meet the required size or structural characteristics are obtained; assigning a unique identifier to each data block and recording its location information in the original document; establishing associations between data blocks, including contextual relationships and hierarchical relationships.

[0026] The vectorization process includes: constructing a training data set, including: taking sentence pairs within the same paragraph as positive sample pairs, randomly selecting sentence pairs from different documents as negative sample pairs, and adding specific prefix tags for different types of data; training an embedding model, including: defining a contrast loss function based on cosine distance, optimizing model parameters using mini-batch stochastic gradient descent, and evaluating model performance using a validation set; generating a vector representation, including: embedding the input text into words, processing the word embedding through a multi-layer Transformer encoder, and obtaining a contextual vector representation of the text; optimizing the vector representation, including: performing vector dimensionality reduction, normalizing the vector, and evaluating the vector quality.

[0027] The vector database adopts a hierarchical index structure, including: the first-level index, including: building an index table containing the basic information of all documents, storing the summary information and keywords of the documents, and establishing the association relationship between documents; the second-level index, including: building a fine-grained index of the specific content of the document, recording the vector representation at the paragraph and sentence level, and storing the semantic relationship between data blocks.

[0028] The retrieval steps include: using multiple similarity calculation methods to calculate the similarity scores between the query vector and all vectors in the vector database; sorting the retrieval results based on the similarity scores; deduplicating the retrieval results to remove duplicate or highly similar content; reordering the retrieval results using a Transformer-based cross encoder; and combining the results of multiple similarity calculation methods to obtain the final ranking.

[0029] Various similarity calculation methods include: Cosine similarity calculation: calculate the similarity through vector dot product and normalize the vector norm; obtain a similarity score in the range of [-1, 1]; Euclidean distance calculation: calculate the sum of the squares of the differences in each dimension of the vector, perform square root operation on the result, and convert the distance into a similarity score; Jaccard similarity calculation: convert the vector into a set representation, calculate the ratio of the set intersection and union, and obtain a similarity score in the range of [0, 1].

[0030] The processing steps of the large language model include: planning the operation path for user queries, including: analyzing the query type and complexity, determining the query decomposition strategy, and selecting a combination of processing methods; decomposing complex queries into multiple sub-queries, including: identifying multiple semantic units in the query, establishing logical relationships between sub-queries, and executing sub-query processing in parallel; processing the context in blocks, including: dividing the retrieval results into blocks according to semantic units, analyzing the semantic associations between blocks, and ensuring contextual coherence; generating multiple candidate answers, including: generating answers based on different context blocks, evaluating the relevance and accuracy of the answers, and selecting the best answer combination; integrating the final response, including: organizing the logical order of multiple answers, ensuring coherence between answers, and generating complete response content.

[0031] Example 2: Based on the method in Example 1, a complete process of providing a multimodal document retrieval enhancement generation method based on a large model is provided: 1. User query input The user inputs the query content through the front-end interface, which can be a question or statement in natural language. The user interface design should be simple and friendly, and support multiple input forms, including text input, voice input, etc. Natural language processing technology (NLP) is used to parse the user input content, extract the core intent and related entities, and ensure the accuracy of subsequent retrieval and generation. The TextProcessor class in the LangChain library is used to pre-process the input text, including removing stop words and standardizing. This process ensures the cleanliness and consistency of the input text, providing a reliable foundation for subsequent embedding generation.

[0032] 2. Embedded model processing Embedded models use the latest large language model technologies, such as BERT and GPT-4, to capture complex semantic information in query content. Multimodal data, including text, images, and audio, is used during model training to improve adaptability to multiple query forms. Use the EmbeddingModel class in the LangChain library to generate query vectors based on the BERT or GPT-4 model. These vectors convert user queries into points in a high-dimensional space, facilitating subsequent similarity calculations and retrieval operations. After the query vector is generated, it is stored in a vector database for subsequent retrieval.

[0033] 3. Multimodal document parsing The parsing module receives multimodal documents in formats including PDF, HTML, TXT, etc. These documents may contain complex layouts, such as pictures, tables, comments, etc. Image recognition technology is used to parse non-text elements such as pictures and tables in the document and convert them into processable structured data. The parsing process combines OCR (optical character recognition) technology to identify and extract text from scanned PDFs and pictures. In addition, natural language processing technology is used to perform grammatical analysis and entity recognition on the text to extract text information and semantic units. The parsed data is divided into text information and semantic units. Figure 1 The overall framework of the parsing module is shown. The text processing function of the LangChain library is used to perform syntax analysis and entity recognition on the text content to extract key information and semantic units.

[0034] To facilitate subsequent segmentation and processing, Figure 2The loading and parsing process of different document formats is demonstrated, including text loader, html loader, pdf loader, word loader, json loader and directory loader. Each loader is optimized for different formats. PDF documents are parsed using PyPDF2 or pdfplumber, HTML pages are parsed using BeautifulSoup4 for DOM, Word documents are processed using python-docx, image files are processed using Pillow and OpenCV, and audio files are feature extracted using librosa to ensure the efficiency and accuracy of the parsing process, and long text information is obtained after parsing.

[0035] 4. Recursive text segmentation Figure 3 The recursive processing flow of text segmentation is shown in detail, and the recursive block strategy is used to segment long text information. The specific steps are as follows: A1: Use the RecursiveTextSplitter class in the LangChain library to recursively divide long text information into smaller data blocks according to predefined delimiters (such as period, comma, semicolon, etc.). The selection of each delimiter is based on the context semantics to ensure that the semantics of the data blocks after segmentation are complete.

[0036] A2: If the first segmentation of the text fails to produce a data block that meets the predetermined size or structural conditions, one or more different delimiters or segmentation criteria are used to recursively reprocess the generated data blocks until a data block that meets the required size or structural characteristics is obtained. This process combines grammatical analysis and semantic understanding to ensure the logical and content coherence of the segmented data blocks.

[0037] A3: Use a single index structure with multiple namespaces or multiple separate index structures to evaluate performance in different data management environments. Use the best index strategy based on data blocks of different types and sizes to improve retrieval efficiency and accuracy. Figure 3 The recursive processing flow of text segmentation is shown. The segmented data blocks are relatively balanced in size and structure, maintaining the semantic integrity and coherence of the content. The segmentation process also includes numbering and marking the data blocks for subsequent retrieval and reorganization.

[0038] Recursive block formula: It's text. is the data block after segmentation, is the size of the predefined block.

[0039] 5. Large language model processing of document content The parsed and segmented data blocks are vectorized using an embedding model based on weakly supervised pre-training and contrastive learning.

[0040] Weakly supervised pre-training uses the intrinsic structure and characteristics of text data as supervisory signals to achieve pre-training in order to learn basic semantic relationships. The key steps are as follows: Paragraph relations as supervisory signals: We use the naturally existing paragraph relations within documents for training. For example, adjacent paragraphs in the same document usually have strong semantic associations, while random paragraphs from different documents usually have low semantic associations. This relationship is used to construct positive and negative sample pairs—adjacent paragraphs as positive samples and paragraphs from different documents as negative samples. By learning these positive and negative relations, the model can gradually grasp the rules of semantic consistency within the document.

[0041] Use of unlabeled data: Taking advantage of weak supervision, the model can be trained on large-scale unlabeled text data without relying on a large amount of manual annotation. This effectively reduces the reliance on high-quality labeled data, while building "pseudo-labels" to simulate labeled data to train preliminary semantic embedding.

[0042] Sentence and paragraph segmentation and embedding: Segment the document and use the segmentation information to train the model so that the model can learn the semantic associations between sentences and paragraphs in context. In addition, use word embedding (such as Word2Vec, GloVe) or sentence embedding (such as BERT) models to initially generate vector representations, and pre-train the model to capture the basic semantics of words, phrases, and paragraphs.

[0043] After weakly supervised pre-training, contrastive learning is used to further improve the ability to capture semantics. The specific steps include: Different types of data are given specific prefixes or labels, such as the prefixes of document paragraphs and titles, so that the model can accurately distinguish content of different semantic categories. This processing effectively generates different vector spaces in contrastive learning, enabling the model to more accurately identify and characterize different types of content.

[0044] Contrastive learning introduces positive and negative sample pairs, which makes the model closer to the distance between positive samples (similar semantics) and farther away from negative samples (different semantics), thereby enhancing the ability to distinguish semantics in a subtle way. This method enables the model to accurately grasp fine-grained semantic differences through detailed similarity learning.

[0045] In the contrastive learning optimization phase, a small-scale, accurately annotated artificial dataset is introduced for fine-tuning the model. The purpose of this annotated dataset is to further optimize the model's embedding representation, making the model more robust and accurate in more complex semantic relationships (such as reasoning relationships and implicit meanings). This step ensures that the model can handle more delicate semantic and emotional expressions.

[0046] After these two stages of training, the embedding model is able to: generate approximate vectors for texts with similar semantics, making similar semantic content more concentrated in the vector space; support a variety of downstream tasks: such as text classification, information retrieval, text clustering, etc. The computability of document semantics after embedding vectorization provides high-quality feature representation for these tasks; based on weak supervision, the model is further optimized through contrastive learning. The model can not only capture the shallow semantic relationships of the text, but also understand more complex and detailed semantic levels, and adapt to the semantic needs in different application scenarios.

[0047] The cosine similarity formula in the vector space model is used to calculate the similarity of texts and measure the semantic proximity between documents or paragraphs. Vector space model formula: Where: :vector, The vector Through this formula, the model can calculate the similarity of each vector, thereby identifying and clustering text fragments with similar semantics.

[0048] The vectorized data blocks are stored in the vector database using a hierarchical index structure: B1: First, build the first index, which contains summary information of all documents, and quickly screen out potentially relevant documents. The summary information uses automatic summary generation technology to extract the most representative sentences or paragraphs from the documents, and uses the FAISS library to manage and search vectors to improve retrieval efficiency.

[0049] B2: Create a second index that covers all specific parts of the document in detail. After initially screening out relevant documents, use the second index to conduct a more detailed and in-depth search. This hierarchical index structure can effectively reduce the complexity of retrieval and improve retrieval speed and accuracy.

[0050] 6. Retrieval and prompt word generation Based on the user query vector, the vector database is searched and the relevant documents, prompt words and query content are returned. The similarity calculation is performed by calculating the similarity between the query vector and the vector in the database, using multiple metrics such as cosine similarity and Euclidean distance. The most similar text paragraph array is selected according to the similarity score sorting. Indexing algorithms (such as IVF_FLAT and HNSW) are used to improve the search efficiency. The vector search and keyword search results are deduplicated and merged to form a new text paragraph array. The FAISS library is used for vector search and index management to ensure the efficiency of the search process. The search results are reordered through a cross encoder, and the Transformer-based cross encoder is used to rearrange the search results and generate a similarity score to ensure that the final results presented are the most relevant and valuable. When the results are displayed, the original metadata of the document is also attached to further increase the credibility of the information.

[0051] Similarity calculation: Calculate the similarity between the user query vector and all vectors stored in the vector database. Similarity calculation uses a variety of metrics, including cosine similarity, Euclidean distance, Jaccard similarity, etc.

[0052] The cosine similarity formula is as follows: The cosine similarity formula is as follows: in, : the norm of vector a, : norm of vector b.

[0053] Euclidean distance formula: :vector, : The first A quantity.

[0054] Jaccard similarity formula: :gather, : The intersection of sets A and B, : The union of sets A and B.

[0055] Retrieval: Sort the results based on the similarity score and select the most similar text paragraph array. Indexing algorithms (such as IVF_FLAT, HNSW, etc.) are also used during the retrieval process to improve retrieval efficiency.

[0056] Deduplication and merging: Deduplication and merging of vector search and keyword search results to form a new text paragraph array. The deduplication process not only considers semantic similarity and text similarity, but also displays the original metadata of the document (such as title, author, creation date, etc.) to increase the credibility of the information and ensure the diversity and relevance of the final results.

[0057] The retrieval results are reordered through a cross encoder, and a Transformer-based cross encoder is used to rearrange the search results and generate similarity scores to ensure that the results presented to users are the most relevant and valuable. When evaluating the relevance of input sentence pairs, the cross encoder can capture deeper semantic relationships and improve the accuracy of retrieval results.

[0058] 7. Large language model generates responses The search results are further processed through the large language model to generate the final response content. The specific steps include: C1: The large language model receives user queries and performs operation path planning after startup, including refining the query content, directly searching for specific data indexes, or combining multiple methods to obtain the best results. In the operation path planning, the model will select the processing method that best suits the current query to ensure that the generated answer is accurate and efficient.

[0059] C2: Use the large language model interface in the LangChain library and GPT-4 to process and generate complex queries. Decompose a single complex query into multiple subqueries, execute each subquery in parallel, collect and merge them into coherent sentences, use them as input data for the large language model, and generate the final answer to the original complex query. The decomposition and execution process of the subquery fully utilizes the advantages of parallel computing and improves processing efficiency. The cross entropy loss formula is used to calculate the loss: : probability distribution, : discrete random variable.

[0060] 8. Response synthesis D1: Send the retrieved context chunks to a large language model, which gradually analyzes and optimizes each part of the context to extract a more precise answer. When processing the context, the model combines the relationship between the contexts to generate a coherent and consistent response.

[0061] D2: Summarize the retrieved context to adapt it to specific prompt conditions, filter and highlight key information, and provide a basis for generating more accurate answers. During the summarization process, natural language processing libraries (such as nltk and spaCy) are used to perform text parsing and entity recognition, automatically extract the most informative parts, reduce the interference of redundant information, and ensure the accuracy of the generated answers.

[0062] D3: Generate multiple targeted answers based on different context blocks, integrate or summarize these answers, and finally form a final reply. During the integration and summarization process, the model will evaluate the quality and relevance of each answer and select the best answer combination to present to the user.

[0063] Through the above detailed technical solutions, the present invention has significant improvements in processing multimodal documents, long text segmentation and information retrieval. Specific innovations include the parsing module processing documents of different formats (such as PDF, HTML, TXT), converting non-text elements (such as pictures, tables) into structured data, and extracting rich semantic information by combining image recognition and natural language processing technology. Pillow and OpenCV are used to process pictures and tables to ensure the comprehensiveness and accuracy of the analysis. Long text processing adopts a recursive block strategy to ensure the accuracy and semantic consistency of text segmentation. The RecursiveTextSplitter class in the LangChain library is used to recursively segment text according to predefined delimiters. Through weakly supervised pre-training and contrastive learning, the vectorization effect of the embedded model is improved, and the hierarchical index structure is combined to achieve fast and accurate document retrieval. The FAISS library is used for efficient vector retrieval and index management. The large language model performs query transformation and response optimization, and the generated answers are more in line with the user's query intention. Through context analysis and summary processing, the response accuracy and relevance are further improved. The large language model interface in the LangChain library is used in combination with GPT-4 to process and generate complex queries.

[0064] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes in form and details may be made without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A multimodal document retrieval enhancement generation method based on a large model, characterized in that: The steps include: S1. receiving a query content input by a user for a multimodal document, parsing the query content input by the user using natural language processing technology, and extracting semantic information; S2. Processing the query content using an embedded model to generate a query vector representing the user's query semantic information, and storing the query vector in a vector database; S3, parsing the multimodal document to obtain long text information, using a recursive block strategy to divide the long text information into data blocks that meet a predetermined size, and numbering the data blocks; S4, using an embedded model to vectorize the data block to generate a document vector and store it in a vector database to build a hierarchical index structure; S4, searching the vector database based on the query vector, and returning a search result; S5. Use a large language model to process the search results to generate response content that meets the user's query intent.

2. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 1 is characterized in that: The multimodal document includes one or more formats of PDF, HTML, TXT, and Markdown; the parsing of the multimodal document includes: using image recognition technology to parse non-text elements in the document and converting them into structured data; using OCR technology to identify and extract text in the document, using natural language processing technology to perform syntax analysis and entity recognition on the extracted text content, and organizing the parsed content into structured data according to a predetermined format.

3. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 2 is characterized in that: The recursive chunking strategy includes: Use predefined delimiters to recursively divide long text information into data blocks; When the first segmentation fails to meet the predetermined conditions, the generated data blocks are reprocessed using different separators or segmentation criteria until data blocks that meet the required size or structural characteristics are obtained; Assign a unique identifier to each data block and record its location information in the original document; Establish associations between data blocks, including contextual relationships and hierarchical relationships.

4. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 3 is characterized in that: The vectorization process includes: Constructing a training dataset, including: taking sentence pairs in the same paragraph as positive sample pairs, randomly selecting sentence pairs from different documents as negative sample pairs, and adding specific prefix tags for different types of data; Train the embedding model, including: defining a contrastive loss function based on cosine distance, optimizing model parameters using mini-batch stochastic gradient descent, and evaluating model performance using a validation set; Generate vector representation, including: embedding the input text into words, processing the word embedding through a multi-layer Transformer encoder, and obtaining the context vector representation of the text; Optimize vector representation, including: perform vector dimensionality reduction, normalize vectors, and evaluate vector quality.

5. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 4 is characterized in that: The vector database adopts a hierarchical index structure, including: The first level of indexing includes: building an index table containing basic information of all documents, storing summary information and keywords of documents, and establishing associations between documents; The second-level index includes: building a fine-grained index of the specific content of the document, recording vector representations at the paragraph and sentence levels, and storing semantic relationships between data blocks.

6. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 5, characterized in that: The retrieval step comprises: A variety of similarity calculation methods are used to calculate the similarity scores between the query vector and all vectors in the vector database; Sort the search results based on similarity scores; De-duplicate the search results to remove duplicate or highly similar content; Use a Transformer-based cross encoder to re-rank the retrieval results; The results of multiple similarity calculation methods are combined to obtain the final ranking.

7. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 6 is characterized in that: The multiple similarity calculation methods include: Cosine similarity calculation: Calculate the similarity through vector dot product and normalize the vector norm; get the similarity score in the range of [-1, 1]; Euclidean distance calculation: Calculate the sum of the squares of the differences in each dimension of the vector, take the square root of the result, and convert the distance into a similarity score; Jaccard similarity calculation: Convert the vector into a set representation, calculate the ratio of the set intersection to the set union, and obtain a similarity score in the range of [0, 1].

8. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 7 is characterized in that: The large language model processing step includes: Plan the operation path for user queries, including: analyzing query type and complexity, determining query decomposition strategy, and selecting a combination of processing methods; Decomposing a complex query into multiple sub-queries, including: identifying multiple semantic units in the query, establishing logical relationships between sub-queries, and executing sub-query processing in parallel; The context is divided into blocks, including: dividing the search results into blocks according to semantic units, analyzing the semantic associations between blocks, and ensuring the coherence of the context; Generate multiple candidate answers, including: generate answers based on different context blocks, evaluate the relevance and accuracy of the answers, and select the best answer combination; Integrate the final response, including organizing the logical order of multiple answers, ensuring coherence between answers, and generating a complete response content.

9. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 8, characterized in that: The following steps are also included: Display the original metadata of the search results, including: document title, author, creation time, document source and version information, and document classification label; Dynamically update the vector database and keyword index according to changes in the data environment, including: monitoring the update status of the data source, detecting new or modified documents, updating the vector representation and index in real time, and maintaining data consistency.

10. The method for enhancing the generation of multimodal document retrieval based on a large model according to claim 9, characterized in that: The receiving method of the user query content includes text input method and voice input method; the text input method supports natural language text input, provides query association and completion functions, and performs input verification in real time; The voice input method uses voice recognition technology to convert voice, reduce noise and enhance signals, and transcribe voice content in real time; Perform natural language processing on the query content, including: performing intent recognition and entity extraction, identifying query keywords and semantic relationships, and building a structured representation of the query.

Citation Information

Cited By

  • Text processing method, electronic equipment, storage medium and program product

    CN120179796A

  • Knowledge construction method and system based on large model and RAG technology

    CN120296111A

  • RAG retrieval generation system oriented to intensive data scene

    CN120296138A

  • Intelligent construction special scheme paragraph intelligent compilation method based on information network

    CN120336334A

  • Biding document multi-mode duplicate checking method and system based on large model

    CN120337898A