Multi-modal document retrieval method and device, electronic equipment and storage medium
Through the multimodal document search method, multimodal documents are preprocessed, parent-child knowledge base is created, and mixed search of dense vectors and sparse vectors is used to solve the shortcomings of existing systems in processing multimodal data and building knowledge bases, and efficient document processing and knowledge base retrieval is achieved.
Patent Information
- Application Number
- CN202510519943.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing search enhancement system is not effective when processing multimodal data (pictures, tables) in documents, the content stored in the knowledge base is incomplete, the relevance of the search results and the problem is low, and it is difficult for non-professionals to build the knowledge base and question-and-answer system independently.
Provide a multimodal document search method, which improves search accuracy and recall by obtaining pending multimodal documents, preprocessing documents to obtain target text lists, creating a parent-child knowledge base, and using a mixed search method of dense vectors and sparse vectors.
Effectively process multimodal data in documents, avoid missing content, build high-quality text libraries and knowledge bases, improve the accuracy and recall of knowledge base retrieval, and simplify the process of non-professionals to build knowledge bases.
Smart Images

Figure CN120030102A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of retrieval enhancement technology, and in particular to a multimodal document retrieval method, device, electronic device and storage medium. Background Art
[0002] With the continuous development of artificial intelligence, large language models and search enhancement generation technology have become key breakthroughs. Among them, search enhancement mainly assists the model in generating answers by retrieving relevant information from external knowledge bases, improving the accuracy and relevance of answers. However, current search enhancement systems generally have similar problems such as difficulty in processing multimodal data (pictures, tables) in documents, incomplete content stored in knowledge bases, low relevance between search results and questions, and difficulty for non-professionals to independently build knowledge bases and question-answering systems. Summary of the invention
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present application provide a multimodal document retrieval method, device, electronic device and storage medium to solve the current retrieval enhancement system. The multimodal data (pictures, tables) in the document is difficult to handle, the content stored in the knowledge base is incomplete, the correlation between the retrieval results and the questions is low, and it is difficult for non-professionals to independently build a knowledge base and question-answering system. Similar problems.
[0004] In order to achieve the above purpose, the technical solutions provided by the embodiments of the present application are as follows: In a first aspect, an embodiment of the present application provides a multimodal document retrieval method, the multimodal document retrieval method comprising: obtaining a multimodal document to be processed; According to the preset regular expression, the multimodal documents to be processed are preprocessed to obtain the target text list. The preprocessing includes: multimodal processing, text cleaning and text segmentation; According to the target text list, a first knowledge base and a second knowledge base are created, there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; According to the dense vector and the sparse vector corresponding to the second knowledge base, the to-be-searched content is searched respectively to obtain a target search result, wherein the target search result is determined according to the sub-segment results respectively retrieved by the dense vector and the sparse vector.
[0005] As an optional implementation, in the first aspect of the embodiment of the present application, the multimodal document to be processed is preprocessed according to a preset regular expression to obtain a target text list, including: According to preset processing rules, multimodal processing is performed on the multimodal document to be processed to obtain a plain text document; Perform text cleaning on plain text documents to obtain the target text content; According to the preset regular expression, the target text content is segmented to obtain the target text list.
[0006] As an optional implementation, in the first aspect of the embodiment of the present application, the target text content is segmented according to a preset regular expression to obtain a target text list, including: According to the title information of the target text content, set the corresponding multi-level regular expression; According to the preset level order, the target text content is matched according to each level of regular expression, and the target text content is segmented according to the matching results to obtain a target text list, which includes the text content obtained after each segmentation.
[0007] As an optional implementation, in the first aspect of the embodiment of the present application, the target text content is matched according to each level of regular expressions in a preset level order, and the target text content is segmented according to the matching results to obtain a target text list, including: According to the preset level order, the target text content is matched according to each level of regular expressions to obtain multiple matching results; When it is detected that the number of characters of the text corresponding to each matching result is within a preset character number range, the target text content is segmented according to the multiple matching results to obtain a target text list.
[0008] As an optional implementation, in the first aspect of the embodiment of the present application, creating a first knowledge base and a second knowledge base according to the target text list includes: The target text list is stored to obtain a first knowledge base; Segment each text content in the target text list to obtain a target sub-text list; The target subtext list is stored to obtain a second knowledge base.
[0009] As an optional implementation, in the first aspect of the embodiment of the present application, the to-be-searched content is searched respectively according to the dense vector and the sparse vector corresponding to the second knowledge base to obtain the target search result, including: Generate dense vectors and sparse vectors based on the target subtext list through the preset embedding model; The content to be searched is searched by using the dense vector and the sparse vector respectively, and the first sub-segment result and the second sub-segment result are obtained; The first sub-segment result and the second sub-segment result are processed to obtain the target search result.
[0010] As an optional implementation, in the first aspect of the embodiment of the present application, the first sub-segment result and the second sub-segment result are processed to obtain the target search result, including: Mapping the first sub-segment result and the second sub-segment result to corresponding text content in the first knowledge base respectively to obtain a plurality of mapped texts; After removing duplicate texts from multiple mapping texts, they are semantically reordered to obtain the target retrieval results.
[0011] In a second aspect, an embodiment of the present application provides a multimodal document retrieval device, the multimodal document retrieval device comprising: an acquisition module, configured to acquire a multimodal document to be processed; A processing module is used to pre-process the multimodal documents to be processed according to a preset regular expression to obtain a target text list, wherein the pre-processing includes: multimodal processing, text cleaning and text segmentation; The processing module is further used to create a first knowledge base and a second knowledge base according to the target text list, there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; The processing module is also used to search the content to be searched according to the dense vector and the sparse vector corresponding to the second knowledge base, and obtain the target search result, wherein the target search result is determined according to the sub-segment results respectively retrieved according to the dense vector and the sparse vector.
[0012] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically used to perform multimodal processing on the multimodal document to be processed according to a preset processing rule to obtain a plain text document; The processing module is specifically used to clean the plain text document to obtain the target text content; The processing module is specifically used to perform text segmentation on the target text content according to a preset regular expression to obtain a target text list.
[0013] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically used to set a corresponding multi-level regular expression according to the title information of the target text content; The processing module is specifically used to match the target text content according to each level of regular expression in a preset level order, and to segment the target text content according to the matching results to obtain a target text list, which includes the text content obtained after each segmentation.
[0014] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically used to match the target text content according to each level of regular expressions in a preset level order to obtain multiple matching results; The processing module is specifically used to segment the target text content according to multiple matching results to obtain a target text list when it is detected that the number of characters of the text corresponding to each matching result is within a preset character number range.
[0015] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically used to store the target text list to obtain a first knowledge base; A processing module, specifically used to segment each text content in the target text list to obtain a target sub-text list; The processing module is specifically used to store the target sub-text list to obtain a second knowledge base.
[0016] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically used to generate a dense vector and a sparse vector according to the target subtext list through a preset embedding model; A processing module, specifically used to search the to-be-searched content through the dense vector and the sparse vector, respectively, to obtain a first sub-segment result and a second sub-segment result; The processing module is specifically used to process the first sub-segment result and the second sub-segment result to obtain the target search result.
[0017] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically used to map the first sub-segment result and the second sub-segment result to corresponding text content in the first knowledge base, respectively, to obtain multiple mapping texts; The processing module is specifically used to remove duplicate texts from multiple mapping texts and perform semantic reordering to obtain target retrieval results.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the multimodal document retrieval method in the first aspect of the embodiment of the present application.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which enables a computer to execute the multimodal document retrieval method in the first aspect of the embodiment of the present application. The computer-readable storage medium includes ROM / RAM, a magnetic disk or an optical disk, etc.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute part or all of the steps of any one of the methods of the first aspect.
[0021] In a sixth aspect, an embodiment of the present application provides an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer executes part or all of the steps of any one of the methods of the first aspect.
[0022] Compared with the prior art, the embodiments of the present application have the following beneficial effects: The embodiments of the present application provide a multimodal document retrieval method, device, electronic device and storage medium, which obtain a multimodal document to be processed; pre-process the multimodal document to be processed according to a preset regular expression to obtain a target text list, wherein the pre-processing includes: multimodal processing, text cleaning and text segmentation; according to the target text list, create a first knowledge base and a second knowledge base, wherein there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; according to the dense vector and the sparse vector corresponding to the second knowledge base, respectively search for the content to be searched to obtain a target search result, wherein the target search result is determined based on the sub-segment results respectively retrieved by the dense vector and the sparse vector. In this solution, the multimodal data (images, tables) in the document are efficiently processed to avoid missing document content. In addition, the document data is cleaned to delete unnecessary data, and regular expressions are used to reasonably segment the document according to the article structure (title, chapter) to build a high-quality text library. A high-quality parent-child knowledge base is built, and a hybrid retrieval method of dense vectors and sparse vectors is used to improve the knowledge base retrieval accuracy and recall rate, and the retrieved text is reordered to obtain the best results, which can effectively improve the processing effect of knowledge documents and the knowledge base retrieval effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 This is a schematic diagram of a multimodal document retrieval method provided in an embodiment of the present application. Figure 1 ; Figure 2 This is a schematic diagram of a multimodal document retrieval method provided in an embodiment of the present application. Figure 2 ; Figure 3It is a schematic diagram of a process of preprocessing a multimodal document to be processed provided by an embodiment of the present application; Figure 4 This is a schematic diagram of a process of knowledge base creation and text retrieval provided in an embodiment of the present application; Figure 5 It is a schematic diagram of a knowledge base evaluation process provided by an embodiment of the present application; Figure 6 It is a structural schematic diagram of a multimodal document retrieval device provided in an embodiment of the present application; Figure 7 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] The terms “first”, “second” and the like in the description and claims of the present application are used to distinguish different objects rather than to describe a specific order of the objects.
[0028] The terms "including" and "having" and any variations thereof in the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0029] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0030] With the continuous development of artificial intelligence, large language models and retrieval enhancement generation technology have become key breakthroughs. Retrieval enhancement generation assists the model in generating answers by retrieving relevant information from an external knowledge base, thereby improving the accuracy and relevance of the answers. This technology can be applied to knowledge question-answering systems to help companies or individuals interpret policy knowledge documents in a certain field. However, there are also some defects in directly using a simple retrieval enhancement generation system, such as: multimodal data (pictures, tables) in documents are difficult to handle, the content stored in the knowledge base is incomplete, the relevance between the retrieval results and the questions is low, the target text in the retrieval results is ranked low, it is difficult for non-professionals to independently build a knowledge base and question-answering system, and it is difficult to evaluate the performance of the knowledge base and question-answering system. The present invention was created in this context and proposes solutions to the above problems.
[0031] The current mainstream knowledge question-answering system is Langchain-Chatchat, which segments text according to a fixed number of words and uses traditional dense vectors to directly retrieve results from the knowledge base. It does not propose further optimization solutions for the above-mentioned problems.
[0032] In order to solve some or all of the above-mentioned technical problems, the embodiments of the present application provide a multimodal document retrieval method, device, electronic device and storage medium to obtain a multimodal document to be processed; pre-process the multimodal document to be processed according to a preset regular expression to obtain a target text list, and the pre-processing includes: multimodal processing, text cleaning and text segmentation; according to the target text list, create a first knowledge base and a second knowledge base, there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; according to the dense vector and sparse vector corresponding to the second knowledge base, search the content to be searched respectively to obtain a target search result, and the target search result is determined according to the sub-segment results retrieved respectively by the dense vector and the sparse vector. In this solution, the multimodal data (images, tables) in the document are efficiently processed to avoid missing document content. In addition, the document data is cleaned to delete unnecessary data, and regular expressions are used to reasonably segment the document according to the article structure (title, chapter) to build a high-quality text library. A high-quality parent-child knowledge base is built, and a hybrid retrieval method of dense vectors and sparse vectors is used to improve the knowledge base retrieval accuracy and recall rate, and the retrieved text is reordered to obtain the best results, which can effectively improve the processing effect of knowledge documents and the knowledge base retrieval effect.
[0033] Embodiment 1: like Figure 1 As shown, Figure 1 A flowchart of a multimodal document retrieval method provided in an embodiment of the present application, the method may include the following steps: 101. Obtain the multimodal document to be processed.
[0034] In an embodiment of the present application, the multimodal document to be processed may be a document containing multiple modal information, and the modal information may include text, images, audio, video, etc. That is to say, the multimodal document to be processed may include data in various forms.
[0035] The multimodal document to be processed involves the processing, fusion and interaction of multiple modal information. The multimodal document to be processed may be a document that a user needs to upload to the knowledge base.
[0036] 102. Preprocess the multimodal document to be processed according to a preset regular expression to obtain a target text list.
[0037] In an embodiment of the present application, the preprocessing may include: multimodal processing, text cleaning and text segmentation.
[0038] It should be noted that since multimodal documents include various forms of data, and the knowledge base may not be able to carry multi-form data due to format restrictions, it is necessary to perform modal processing on the multimodal documents, that is, convert all non-text data into text data, so as to ensure that the data format in the knowledge base is consistent; in addition, there may be duplicate data or unnecessary data in the document, so it is necessary to perform text cleaning on the document to ensure that the processed text data is normal and useful data; and, the document includes a large amount of text information. If all of it is stored directly in the knowledge base, then the storage content is large, and there is no classification, no chapter division and other information, which may make the retrieval more complicated and inefficient, so it is necessary to segment the text, which can be based on title chapter information, or paragraph information, etc. Specifically, text segmentation can be achieved according to preset regular expressions.
[0039] It should be noted that regular expression (regex / regexp) is a pattern description language used for text search, replacement, parsing and other operations. It uses a specific syntax to define a search pattern, which describes the rules of character sequences in the text. Regular expressions are very important in computer science, especially in text processing, data cleaning, data validation and other fields. For text segmentation, the re.split() function is very useful. This function will split the string according to the content matched by the regular expression and return the segmented list.
[0040] In an embodiment of the present application, the target text list obtained after preprocessing can be understood as a collection of multiple text contents, each of which is text information obtained after the multimodal document to be processed is processed through multimodal processing, text cleaning and text segmentation.
[0041] 103. Create a first knowledge base and a second knowledge base according to the target text list.
[0042] In the embodiment of the present application, after obtaining the target text list, a knowledge base can be created according to the text content in the target text list.
[0043] It should be noted that, generally speaking, the length of the questions that users want to search for answers to will not be too long (for example, within 30 words), and the characteristic of vector search is that it will give priority to returning results that are close to the length of the question. In other words, the closer the length of the text content that can be searched and stored in the knowledge base is to the length of the question (the shorter it is), the easier it is for the vector search to return the target result; therefore, in order to improve the knowledge base search effect, the first knowledge base and the second knowledge base can be constructed according to the target text list, which can also be called parent-child knowledge bases. From the name, it can be understood that the child knowledge base is usually a detailed expansion and deepening of one or some knowledge points in the parent knowledge base, providing more specific and detailed information.
[0044] In the embodiment of the present application, there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base. It can be understood that the first knowledge base is the parent knowledge base and the second knowledge base is the child knowledge base, that is, after the first knowledge base is created, more sub-contents are obtained after further text segmentation, thereby forming the second knowledge base, and when the text content in the first knowledge base is segmented, the sub-contents obtained by segmentation will generate an index relationship, which is convenient for text tracing.
[0045] 104. According to the dense vector and the sparse vector corresponding to the second knowledge base, the contents to be searched are searched respectively to obtain target search results.
[0046] In an embodiment of the present application, since the second knowledge base is obtained by further segmenting the first knowledge base, it can be understood that the text content in the second knowledge base is more detailed, so dense vectors and sparse vectors can be generated for the text content of the second knowledge base, wherein the dense vector tends to favor the semantic similarity between the question and the text during retrieval, while the sparse vector pays more attention to the matching degree of key words. Therefore, a mixed retrieval mode of dense vectors and sparse vectors is used to perform a mixed retrieval on the content to be retrieved to obtain a target retrieval result. The content to be retrieved can be a question input by a user that needs to be retrieved and answered, and the target retrieval result is determined based on the sub-segment results retrieved by the dense vectors and the sparse vectors respectively.
[0047] A multimodal document retrieval method provided in an embodiment of the present application can efficiently process the multimodal data (images, tables) content in a document to avoid missing document content. In addition, the document data is cleaned to delete unnecessary data, and regular expressions are used to reasonably segment the document according to the article structure (title, chapter) to build a high-quality text library. A high-quality parent-child knowledge base is built, and a hybrid retrieval method of dense vectors and sparse vectors is used to improve the knowledge base retrieval accuracy and recall rate, and the retrieval text is reordered to obtain the best results, which can effectively improve the processing effect of knowledge documents and the knowledge base retrieval effect.
[0048] like Figure 2 As shown, Figure 2 A flowchart of a multimodal document retrieval method provided in an embodiment of the present application, the method may also include the following steps: 201. Obtain a multimodal document to be processed.
[0049] In the embodiment of the present application, for the description of step 201, please refer to the detailed description of step 101 in the above embodiment, and the embodiment of the present application will not be repeated.
[0050] 202. Perform multimodal processing on the multimodal document to be processed according to preset processing rules to obtain a plain text document.
[0051] In the embodiment of the present application, a multimodal large language model and a universal large language model are used to process non-text content such as pictures and tables in the multimodal document to be processed, and they are processed into text information and mapped back to the original document. The preset processing rules can be pre-set processing methods for various non-text content, and the rules can be set by the user.
[0052] In some embodiments, the document content can be traversed first, and no processing is performed when plain text content is encountered; when a table is traversed, the table can be converted into a markdown format that can be understood by a general large language model, and then the large language model is allowed to summarize the content in the markdown table to obtain a detailed summary text, and then the markdown table and the summary text are combined into a whole, and [table] and [ / table] tags are added at the beginning and end respectively, and finally this part of the text is used to replace the original table in the document; when a picture is traversed, the picture can be saved as a specified file directory, and then the multimodal large language model is used to analyze the content in the picture to generate a summary text, and then [img] and [ / img] tags are added at the beginning and end of the text respectively, and a tag corresponding to the original picture is added at the end, and finally this part of the text is used to replace the original picture in the document.
[0053] 203. Perform text cleaning on the plain text document to obtain the target text content.
[0054] In the embodiments of the present application, a document can generally be divided into three parts: the beginning, the main text and the appendix. When constructing a knowledge base, only the main text needs to be extracted, so the beginning content and the appendix content can be text cleaned; in addition, repeated content in plain text documents can also be cleaned, for example: after multimodal processing of the table content, there may be repeated explanatory content, and the repeated content can be deleted.
[0055] 204. Perform text segmentation on the target text content according to a preset regular expression to obtain a target text list.
[0056] It should be noted that if Figure 3 The figure shows a flowchart for preprocessing the multimodal documents to be processed. It can be seen that the plain text content, image content and table content in the multimodal documents uploaded by the user are processed differently to obtain a multimodal processed document, which is a plain text document. In addition, the user can choose to preset data or design regular expressions independently, and perform text processing on the multimodal processed document based on regular expressions. After the text is cleaned to remove redundant data, the text is segmented according to the article structure (title) to obtain a target text list.
[0057] In some embodiments, a corresponding multi-level regular expression is set according to the title information of the target text content; the target text content is matched according to each level of regular expression in a preset level order, and the target text content is segmented according to the matching results to obtain a target text list, which includes the text content obtained after each segmentation.
[0058] It should be noted that, for most knowledge documents with standard formats, the titles of the chapters in the main text are relatively uniform and formatted. The first-level titles are mostly named in Chinese numerals / Arabic numerals + title name (such as "I.XX", "1.XX") or Chapter X, the second-level titles are mostly named in brackets + Chinese numerals + title name (such as "(I)XX") or two-level Arabic numerals + title name (such as "1.1 XX") or Section X / Article, and the third-level titles are mostly named in brackets + Arabic numerals + title name (such as "(1)XX") or three-level Arabic numerals + title name (such as "1.1.2 XX"). Therefore, according to the above title information, multiple regular expression matching templates for the first, second, and third-level titles can be preset. The preset level order can be from low to high, that is, starting text segmentation from the lowest level title, traversing the entire document in turn, and obtaining the target text list.
[0059] In some embodiments, assuming that the third-level title is set as the minimum matching unit for document division, whenever a first-level or second-level title is matched, the first-level and second-level titles can be temporarily stored first. After matching the third-level title, all the contents under the third-level title can be read until the next first-, second-, or third-level title is matched, and then the previously temporarily stored first-, second-, or third-level titles and the contents under the third-level titles are spliced into a text segment and stored in the result text list; if the document does not have a third-level title, the second-level title is used as the minimum matching unit for document division by default. If there is no second-level title, the first-level title is used as the minimum matching unit.
[0060] In some embodiments, in order to prevent the total number of characters in each text segment from being too small, or the length of each text segment from being too long, a preset character range can be set. Specifically, the target text content is matched according to each level of regular expression in a preset level order to obtain multiple matching results; when it is detected that the number of characters of the text corresponding to each matching result is within the preset character range, the target text content is segmented according to the multiple matching results to obtain a target text list.
[0061] It should be noted that in order to prevent the total number of characters in each text segment from being too small, whenever the total number of characters of all third-level titles and their contents under a second-level title is less than the lower limit threshold of a preset character number range (for example, 300), all the contents under the second-level title can be regarded as a text segment and stored in the result text list; similarly, whenever the total number of characters of all second- and third-level titles and their contents under a first-level title is less than the lower limit threshold, all the contents under the first-level title can be regarded as a text segment and stored in the result text list; finally, the result text list can be traversed, and all text segments with less than the lower limit threshold are pieced together with nearby text segments until the number of characters exceeds the lower limit threshold.
[0062] It should be noted that, similarly, in order to prevent each text segment from being too long, when traversing the target text list, each text segment that exceeds the upper threshold of the preset character number range (e.g., 800) is split twice according to the preset chunk_size (e.g., 500) and over_lap (e.g., 150). The number of characters of each text content in the target text list thus obtained is within the preset character number range.
[0063] In some embodiments, if a text segment contains summary text content of a table or image ([table] or [img]) obtained through multimodal processing, in order to ensure the integrity and readability of the text, each [table] or [img] content will be completely segmented according to the [table] or [img] tag to avoid incomplete [table] or [img] content after segmentation.
[0064] 205. Store the target text list to obtain a first knowledge base.
[0065] In an embodiment of the present application, the target text list may be stored in a sqlite database as a first knowledge base.
[0066] 206. Segment each text content in the target text list to obtain a target sub-text list.
[0067] In an embodiment of the present application, each text content in the target text list can be regarded as a parent text, and then each parent text is divided twice according to a preset chunk_size (for example, 100). Each parent text is divided into multiple child texts, and the multiple child texts corresponding to all parent texts are collected to obtain a target child text list.
[0068] It should be noted that each parent text will be divided into multiple child texts, and then an index mapping relationship needs to be established between the parent text and the multiple child texts.
[0069] 207. Store the target subtext list to obtain a second knowledge base.
[0070] In an embodiment of the present application, the target subtext list may be stored in a milvus vector library as a second knowledge base.
[0071] 208. Generate dense vectors and sparse vectors according to the target sub-text list through the preset embedding model.
[0072] In an embodiment of the present application, when storing the target sub-text list into a vector library, a preset embedding model can also be used to simultaneously generate dense vectors and sparse vectors for the target sub-text list. The preset embedding model can be an embedding embedding model.
[0073] In some embodiments, dense vectors are typically generated through deep learning models (such as Word2Vec, BERT, etc.), while sparse vectors may be generated through some feature engineering methods (such as bag-of-words model, TF-IDF, etc.).
[0074] 209. The content to be searched is searched by using the dense vector and the sparse vector respectively to obtain a first sub-segment result and a second sub-segment result.
[0075] In an embodiment of the present application, when searching the knowledge base, a dense + sparse hybrid retrieval method is used, and the content to be searched can be brought into a dense vector for retrieval to obtain the first sub-segment result, and then the content to be searched can be brought into a sparse vector for retrieval to obtain the second sub-segment result.
[0076] It should be noted that the first sub-segment result and the second sub-segment result can be considered as the content in the dense vector and the sparse vector that matches the content to be retrieved.
[0077] It should be noted that there is no limit to the number of the first sub-segment results and the second sub-segment results, and there can be multiple results. The number can be a preset fixed value or a non-fixed value. For example, the top_k (for example, 50) sub-segment results are returned from the dense vector and the sparse vector respectively. The combination of the two vector retrievals can more efficiently retrieve results from the knowledge base and improve the recall rate.
[0078] 210. Process the first sub-segment result and the second sub-segment result to obtain a target search result.
[0079] In some embodiments, specifically, the first sub-segment result and the second sub-segment result may be mapped to corresponding text content in the first knowledge base respectively to obtain multiple mapped texts; after deleting duplicate texts from the multiple mapped texts, semantically reordering is performed to obtain the target retrieval result.
[0080] It should be noted that since the second knowledge base can be considered as a child knowledge base, and there is an index mapping relationship between the second knowledge base and the second knowledge base, that is, the text content in the second knowledge base has corresponding text in the first knowledge base, so the first sub-segment result and the second sub-segment result can be mapped back to the corresponding parent text, and duplicate data can be deleted. Then, the rerank model is used for semantic reordering to finally obtain the target retrieval result. The target retrieval result can be multiple results, and the number can be a pre-set fixed value or a non-fixed value. The target retrieval result can be the text content that best matches the content to be retrieved.
[0081] It should be noted that if Figure 4 The figure shows a flow chart of knowledge base creation and text retrieval. It can be seen that the target text list is stored in the SQLite database to obtain the parent knowledge base, and then the secondary segmentation is performed to establish a mapping to obtain sub-segment data. After vectorization through the embedding model, dense vectors and sparse vectors are generated to obtain a sub-knowledge base. After the user enters the question to be retrieved and answered, multiple sub-segment results are obtained through the dense + sparse hybrid retrieval method. After mapping the parent segment, duplicates are removed to obtain the parent segment result, which is re-ranked through the rerank model to obtain the final target retrieval result. In addition, a large language model can be combined for retrieval and answering.
[0082] A multimodal document retrieval method provided in an embodiment of the present application can efficiently process the multimodal data (images, tables) content in a document to avoid missing document content. In addition, the document data is cleaned to delete unnecessary data, and regular expressions are used to reasonably segment the document according to the article structure (title, chapter) to build a high-quality text library. A high-quality parent-child knowledge base is built, and a hybrid retrieval method of dense vectors and sparse vectors is used to improve the knowledge base retrieval accuracy and recall rate, and the retrieval text is reordered to obtain the best results, which can effectively improve the processing effect of knowledge documents and the knowledge base retrieval effect.
[0083] In some embodiments, according to the above method, a retrieval enhancement generation system can be constructed, which integrates functions such as user registration, knowledge base creation, document upload, multimodal data processing, document preprocessing, knowledge base addition, deletion, modification and query, large language model question and answer, and retrieval enhancement generation (RAG) system evaluation. The RAG full-process framework is opened up, so that non-professionals can pre-process multimodal documents according to the instruction manual, efficiently create domain knowledge bases and question and answer systems, evaluate knowledge bases and systems, and can update and modify the knowledge base in real time.
[0084] The system can be divided into four modules: user registration (Register), knowledge base (KnowledgeBase), conversation (Chat) and evaluation (Evaluation).
[0085] User registration module: Every new user needs to register an account in the system and can only use it officially after successful registration.
[0086] Knowledge base module: divided into knowledge base creation, document upload, knowledge base addition and deletion, document preprocessing status query, knowledge base status query, knowledge base deletion and other parts. Specifically, users need to create a knowledge base first, and then upload the documents to be added to the knowledge base. When uploading documents, users can freely choose whether to perform multimodal processing on the documents (parse chart data) and whether to split the documents according to the article structure (chapter, title). If users choose to split the documents according to the article structure (chapter, title), they need to set the regular expressions of the titles at all levels in the document at the same time. The system will preset regular expressions of various types of titles for users, and users can choose according to their needs. After the document is uploaded, the system background will automatically start multimodal preprocessing and document preprocessing. Users can obtain the current document processing process information through the preprocessing status query function. After the document preprocessing is successful, the system will automatically start to build the parent-child knowledge base. Users can obtain the document information that has been successfully added to the knowledge base through the knowledge base status query function. If users want to update the knowledge base content, they can update it by uploading new documents or deleting specified documents. At the same time, users can also delete unwanted knowledge bases through the knowledge base deletion function. In addition, the system also has an exception handling function. If any module such as document upload, preprocessing, knowledge base addition and deletion fails, users can obtain fault information through the status query function.
[0087] Dialogue module: Users can search the knowledge base after the document is successfully added to the knowledge base. Specifically, users need to specify the target knowledge base, and can independently set parameters such as the number of search results returned (top_n), vector search threshold (threshold), etc. In addition, users can also choose whether to configure the large language model Q&A. If selected, the system will not only return the knowledge base search results, but also return the answers to the user's questions based on the search results by the large language model; in the large language model Q&A function, the system will preset the system prompt word template and the user prompt word template, but users can also customize the system prompt words according to the field to which the current knowledge base belongs. For example, when the knowledge base belongs to the field of environmental protection, you can design prompt words like "You are an environmental expert who is familiar with professional knowledge in environmental protection" to give the large model a specific personality, so that it can answer user questions from a professional perspective in a more rigorous and formal way.
[0088] Evaluation module: Users can evaluate the performance of a specified knowledge base and its question-answering system by uploading a local test data set. This system integrates the Ragas (Retrieval Augmented Generation Evaluation) framework, including test set upload and status query functions. Specifically, users need to upload test data, and can only upload test sets in json format, which can be question-answer pairs: [{"question1":"","answer1":""}, {"question2":"","answer2":""},...]; users can also choose whether to configure a large language model. If not, the system will only evaluate the retrieval recall and precision of the knowledge base; if configured, the system will simultaneously evaluate the knowledge base retrieval performance and the accuracy of the large language model's answers (i.e., the performance of the question-answering system); users can obtain evaluation progress information through the status query function.
[0089] In some embodiments, Figure 5 The figure shows the specific flow chart of knowledge base evaluation. The user uploads the evaluation data set, and the evaluation results are obtained according to the Ragas evaluation method through the child knowledge base, parent knowledge base, LLM model, rerank model and embedding model.
[0090] This method is based on Django technology to build a retrieval enhancement generation system, which integrates functions such as user registration, knowledge base creation, document upload, multimodal data processing, document preprocessing, knowledge base addition, deletion, modification and query, large language model question and answer, and RAG system evaluation. It opens up the whole process framework of retrieval enhancement generation (RAG), so that enterprise users who are not computer professionals can also preprocess multimodal documents according to the instruction manual, efficiently create domain knowledge base and question and answer system, evaluate the knowledge base and system, and update and modify the knowledge base in real time.
[0091] Embodiment 2: like Figure 6 As shown, an embodiment of the present application provides a multimodal document retrieval device, which may include: An acquisition module 601 is used to acquire a multimodal document to be processed; The processing module 602 is used to pre-process the multimodal document to be processed according to a preset regular expression to obtain a target text list, wherein the pre-processing includes: multimodal processing, text cleaning and text segmentation; The processing module 602 is further used to create a first knowledge base and a second knowledge base according to the target text list, wherein there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; The processing module 602 is also used to search the content to be searched according to the dense vector and the sparse vector corresponding to the second knowledge base to obtain a target search result, which is determined according to the sub-segment results obtained by searching the dense vector and the sparse vector respectively.
[0092] In some embodiments, the processing module 602 is specifically used to perform multimodal processing on the multimodal document to be processed according to a preset processing rule to obtain a plain text document; Processing module 602 is specifically used to perform text cleaning on the plain text document to obtain the target text content; The processing module 602 is specifically used to perform text segmentation on the target text content according to a preset regular expression to obtain a target text list.
[0093] In some embodiments, the processing module 602 is specifically used to set a corresponding multi-level regular expression according to the title information of the target text content; Processing module 602 is specifically used to match the target text content according to each level of regular expression in a preset level order, and segment the target text content according to the matching results to obtain a target text list, which includes the text content obtained after each segmentation.
[0094] In some embodiments, the processing module 602 is specifically used to match the target text content according to each level of regular expressions in a preset level order to obtain multiple matching results; The processing module 602 is specifically configured to segment the target text content according to the multiple matching results to obtain a target text list when it is detected that the number of characters of the text corresponding to each matching result is within a preset character number range.
[0095] In some embodiments, the processing module 602 is specifically used to store the target text list to obtain a first knowledge base; Processing module 602 is specifically used to segment each text content in the target text list to obtain a target sub-text list; The processing module 602 is specifically used to store the target sub-text list to obtain a second knowledge base.
[0096] In some embodiments, the processing module 602 is specifically configured to generate a dense vector and a sparse vector according to the target subtext list by using a preset embedding model; The processing module 602 is specifically used to search the content to be searched by using the dense vector and the sparse vector respectively to obtain the first sub-segment result and the second sub-segment result; The processing module 602 is specifically configured to process the first sub-segment result and the second sub-segment result to obtain a target search result.
[0097] In some embodiments, the processing module 602 is specifically configured to map the first sub-segment result and the second sub-segment result to corresponding text content in the first knowledge base, respectively, to obtain a plurality of mapped texts; The processing module 602 is specifically used to remove duplicate texts from the multiple mapping texts and perform semantic reordering to obtain a target search result.
[0098] In the embodiments of the present application, each module can implement the multimodal document retrieval method provided by the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0099] Embodiment three: like Figure 7 As shown, an embodiment of the present application further provides an electronic device, which may include: A memory 701 storing executable program codes; a processor 702 coupled to the memory 701; The processor 702 calls the executable program code stored in the memory 701 to execute the multimodal document retrieval method executed by the electronic device in the above-mentioned method embodiments.
[0100] Embodiment 4: An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the multimodal document retrieval method in the above method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0101] The embodiment of the present application also provides a computer program product, which stores a computer program. When the computer program is executed by a processor, each process of the multimodal document retrieval method in the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0102] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media that include computer-usable program code.
[0103] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a module, a program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0104] In the present application, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0105] In this application, memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0106] In this application, a person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, and the computer-readable medium includes permanent and non-permanent, removable and non-removable storage media. The storage medium can implement information storage by any method or technology, and the information can be a computer-readable instruction, a data structure, a module of a program, or other data. Examples of computer storage media include, but are not limited to, Parallel Random Access Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Programmable Read-only Memory (PROM), Erasable Programmable Read Only Memory (EPROM), other types of Random Access Memory (RAM), Read-Only Memory (ROM), One-time Programmable Read-Only Memory (OTPROM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), flash memory or other memory technology, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this article, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0107] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0108] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application. The above-mentioned multiple embodiments are not necessarily multiple independent embodiments. Dividing into multiple embodiments is only used to highlight the different technical features in different embodiments. Those skilled in the art should be aware that the above-mentioned multiple embodiments can also be combined in any way.
[0109] In the various embodiments of the present application, it should be understood that the size of the serial numbers of the above-mentioned processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0110] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0112] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a memory and includes several requests for a computer device (which can be a personal computer, server or network device, etc., specifically a processor in a computer device) to perform some or all of the steps of the above-mentioned methods of various embodiments of the present application.
[0113] The above is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal document retrieval method, characterized in that: The method comprises: Get the multimodal document to be processed; Preprocessing the multimodal document to be processed according to a preset regular expression to obtain a target text list, wherein the preprocessing includes: multimodal processing, text cleaning and text segmentation; According to the target text list, a first knowledge base and a second knowledge base are created, wherein there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; According to the dense vector and the sparse vector corresponding to the second knowledge base, the to-be-searched content is searched respectively to obtain a target search result, wherein the target search result is determined according to the sub-segment results respectively searched by the dense vector and the sparse vector.
2. The method according to claim 1, characterized in that The method of preprocessing the multimodal document to be processed according to a preset regular expression to obtain a target text list includes: According to preset processing rules, multimodal processing is performed on the multimodal document to be processed to obtain a plain text document; Performing text cleaning on the plain text document to obtain target text content; The target text content is segmented according to the preset regular expression to obtain the target text list.
3. The method according to claim 2, characterized in that The step of performing text segmentation on the target text content according to the preset regular expression to obtain the target text list includes: According to the title information of the target text content, a corresponding multi-level regular expression is set; According to the preset level order, the target text content is matched according to each level of regular expression, and the target text content is segmented according to the matching results to obtain the target text list, which includes the text content obtained after each segmentation.
4. The method according to claim 3, characterized in that The target text content is matched according to each level of regular expressions in a preset level order, and the target text content is segmented according to the matching results to obtain the target text list, including: According to the preset level order, the target text content is matched according to each level of regular expressions to obtain multiple matching results; When it is detected that the number of characters of the text corresponding to each matching result is within a preset character number range, the target text content is segmented according to the multiple matching results to obtain the target text list.
5. The method according to claim 1, characterized in that The step of creating a first knowledge base and a second knowledge base according to the target text list includes: Storing the target text list to obtain the first knowledge base; Segment each text content in the target text list to obtain a target sub-text list; The target subtext list is stored to obtain the second knowledge base.
6. The method according to claim 5, characterized in that The step of searching the content to be searched according to the dense vector and the sparse vector corresponding to the second knowledge base to obtain the target search result includes: Generate the dense vector and the sparse vector according to the target subtext list by using a preset embedding model; Retrieving the content to be searched by using the dense vector and the sparse vector respectively to obtain a first sub-segment result and a second sub-segment result; The first sub-segment result and the second sub-segment result are processed to obtain the target search result.
7. The method according to claim 6, characterized in that The processing of the first sub-segment result and the second sub-segment result to obtain the target search result includes: Mapping the first sub-segment result and the second sub-segment result to corresponding text contents in the first knowledge base respectively to obtain a plurality of mapped texts; The multiple mapping texts are deleted and semantically reordered to obtain the target retrieval result.
8. A multimodal document retrieval device, characterized in that: include: An acquisition module, used to acquire multimodal documents to be processed; A processing module, used for preprocessing the multimodal document to be processed according to a preset regular expression to obtain a target text list, wherein the preprocessing includes: multimodal processing, text cleaning and text segmentation; The processing module is further used to create a first knowledge base and a second knowledge base according to the target text list, wherein there is an index relationship between the first knowledge base and the second knowledge base, and the second knowledge base is obtained by segmenting the first knowledge base; The processing module is also used to search the content to be searched according to the dense vector and the sparse vector corresponding to the second knowledge base, respectively, to obtain a target search result, wherein the target search result is determined based on the sub-segment results respectively retrieved according to the dense vector and the sparse vector.
9. An electronic device, characterized in that: include: A memory storing executable program code; and a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the multimodal document retrieval method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the multimodal document retrieval method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Nuclear power knowledge base retrieval method and system fusing semantic features and keywords
CN118626591A
RAG-based vertical domain knowledge multi-round question and answer method
CN118964556A
Power system vector knowledge base construction method based on text vectorization
CN118964695A
Question and answer method and system based on paragraph splitting and knowledge engine platform
CN119271779A
Multi-modal knowledge base management method and device and storage medium
CN119272756A
Cited By
Intelligent question answering method and device based on retrieval enhancement technology and electronic equipment
CN120832400A
Text data governance method based on large model
CN121503443A
Railway standard intelligent question-answering method and system fusing standard features and retrieval enhancement
CN122388124A