Retrieval enhancement processing method and system based on large model
Through adaptive analysis and segmentation of knowledge documents, combined with multi-routing strategies of sparse search and dense search, semantic correlation scoring is used to use the fine-scheduling model to solve the problem of insufficient knowledge retrieval recall in the RAG question-and-answer system, and the accuracy of knowledge question-and-answer and intelligent dialogue is improved.
Patent Information
- Application Number
- CN202510719249.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Due to insufficient knowledge retrieval recall, the existing RAG question and answer system has low accuracy in the content of knowledge question and answer and intelligent dialogue. Especially when user questions are expressed one-sidedly or involve knowledge details, it is impossible to recall relevant knowledge points in a comprehensive and multi-faceted manner.
The search enhancement processing method based on large models is adopted, and knowledge documents are adaptively analyzed and segmented, combined with the multi-routing retrieval recall strategy of sparse search and dense search, and the knowledge recall results are refined using the fine-scheduling model, and semantic correlation scores are performed to finally determine the search results.
It improves the accuracy and recall rate of knowledge retrieval, enhances the accuracy of knowledge Q&A and intelligent dialogue content of the RAG Q&A system, especially when processing knowledge documents of different formats and types, significantly improves the search effect.
Smart Images

Figure CN120256648A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a retrieval enhancement processing method and system based on a large model. Background Art
[0002] A Retrieval Augmented Generation (RAG) question-and-answer system is a knowledge-enhanced question-and-answer system that mainly answers information inquiries in specific knowledge fields or knowledge-based questions in open fields.
[0003] In existing RAG question-and-answer systems, since knowledge retrieval recall usually uses semantic retrieval based on cosine similarity in a database and text content retrieval, etc., in cases where the user's question expression is relatively one-sided or the question itself belongs to knowledge details, etc., simple knowledge text matching cannot comprehensively and multi-facetedly recall the knowledge block where the corresponding knowledge point of the question is located, resulting in the failure of knowledge retrieval recall in the RAG question-and-answer system, thereby reducing the accuracy of knowledge answering and intelligent dialogue content in the RAG question-and-answer system.
[0004] Therefore, how to improve the knowledge retrieval of the RAG question-and-answer system to improve the accuracy of knowledge answering and intelligent dialogue content in the RAG question-and-answer system is an urgent problem to be solved in this application. Summary of the Invention
[0005] In view of this, this application discloses a retrieval enhancement processing method and system based on a large model, aiming to improve the knowledge retrieval accuracy of the retrieval augmented generation question-and-answer system, thereby improving the accuracy of knowledge answering and intelligent dialogue content in the retrieval augmented generation question-and-answer system.
[0006] To achieve the above object, the disclosed technical solution is as follows:
[0007] The first aspect of this application discloses a retrieval enhancement processing method based on a large model, and the method includes:
[0008] When receiving a user question, obtain the knowledge document corresponding to the user question;
[0009] Perform adaptive parsing and segmentation on the knowledge document to obtain a segmentation result;
[0010] Based on a multi-routing retrieval recall strategy and the user question, perform retrieval recall on the segmentation result to obtain a knowledge recall result; wherein, the multi-routing retrieval recall strategy is determined by sparse retrieval and dense retrieval;
[0011] Based on a re-ranking model, perform re-ranking on the knowledge recall result to obtain a re-ranking result;
[0012] Sort the refined ranking results according to a preset sorting method to determine the retrieval results.
[0013] Preferably, the adaptive parsing and segmentation of the knowledge document to obtain a segmentation result includes:
[0014] Identify each feature type in the knowledge document;
[0015] According to each feature type, call the corresponding parsing method to perform content parsing of the adaptive knowledge corpus on the knowledge document to obtain a parsing result;
[0016] Perform corpus segmentation on the parsing result by combining a hierarchical structure and linear segmentation by character length to obtain a segmentation result.
[0017] Preferably, the performing corpus segmentation on the parsing result by combining a hierarchical structure and linear segmentation by character length to obtain a segmentation result includes:
[0018] Perform text segmentation on the parsing result to obtain the content of each chapter; wherein, the content of each chapter includes at least the chapter length and the chapter text content;
[0019] Determine the length of each chapter from the content of each chapter;
[0020] If the consecutive chapter lengths at the same level in the lengths of each chapter are less than or equal to a preset threshold, merge the consecutive chapter contents corresponding to the consecutive chapter lengths at the same level;
[0021] If the length of a single chapter in the lengths of each chapter is less than or equal to the preset threshold, merge the consecutive chapter contents at the same level;
[0022] If the length of a single chapter in the lengths of each chapter is greater than the preset threshold, perform a threshold length judgment on the next-level directory of the current-level chapter until the length of a single chapter in the last-level directory is greater than the preset threshold, and complete the threshold segmentation according to the line break characters and punctuation marks of the chapter content to complete the corpus segmentation of the parsing result and obtain a segmentation result.
[0023] Preferably, the retrieving and recalling the segmentation result based on a multi-routing retrieval recall strategy and the user question to obtain a knowledge recall result includes:
[0024] Set a global cosine similarity threshold for knowledge search and the number of retrieved knowledge chunks;
[0025] Extract the keywords of the user question;
[0026] The embedding model is used to perform vector representations on the user question and the keyword respectively, obtaining the sparse vector of the retrieval query, the dense vector of the retrieval query, the sparse vector of the retrieval keyword, and the dense vector of the retrieval keyword;
[0027] The open-source vector database is used to perform sparse vector retrieval on the sparse vector of the retrieval query to obtain the knowledge block corresponding to the first sparse vector, perform dense vector retrieval on the dense vector of the retrieval query to obtain the knowledge block corresponding to the first dense vector, and de-duplicate and merge the knowledge block corresponding to the first sparse vector and the knowledge block corresponding to the first dense vector to obtain the retrieval query matching knowledge block;
[0028] The open-source vector database is used to perform sparse vector retrieval on the sparse vector of the retrieval keyword to obtain the knowledge block corresponding to the second sparse vector, perform dense vector retrieval on the dense vector of the retrieval keyword to obtain the knowledge block corresponding to the second dense vector, and de-duplicate and merge the knowledge block corresponding to the second sparse vector and the knowledge block corresponding to the second dense vector to obtain the retrieval keyword matching knowledge block;
[0029] The open-source vector database is used to perform sparse vector retrieval for knowledge generation question matching on the sparse vector of the retrieval query to obtain the knowledge block corresponding to the third sparse vector, perform dense vector retrieval for knowledge generation question matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the third dense vector, and merge the knowledge block corresponding to the third sparse vector and the knowledge block corresponding to the third dense vector to obtain the knowledge generation question matching knowledge block;
[0030] The open-source vector database is used to perform sparse vector retrieval for abstract matching on the sparse vector of the retrieval query to obtain the knowledge block corresponding to the fourth sparse vector, perform dense vector retrieval for abstract matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the fourth dense vector, and de-duplicate and merge the knowledge block corresponding to the fourth sparse vector and the knowledge block corresponding to the fourth dense vector to obtain the abstract matching knowledge block;
[0031] The multi-route retrieval result formula is used to calculate the scores of the retrieval query matching knowledge block, the retrieval keyword matching knowledge block, the knowledge generation question matching knowledge block, and the abstract matching knowledge block, obtaining the comprehensive score of all retrieval results;
[0032] The comprehensive score of all retrieval results is filtered by the number of retrieved knowledge blocks and the global cosine similarity threshold of knowledge search to obtain the knowledge recall result.
[0033] Preferably, the knowledge recall result is refined by the refinement model to obtain the refinement result, including:
[0034] The refined ranking model fine-tuned with data in the communication field scores the semantic relevance between the knowledge chunks in the knowledge recall results and the user questions to obtain respective semantic relevance scores, so as to complete the process of refining the knowledge recall results based on the refined ranking model to obtain the refined ranking results.
[0035] Preferably, sorting the refined ranking results according to a preset sorting method to determine the retrieval results includes:
[0036] Sorting the respective semantic relevance scores in the refined ranking results according to a preset sorting method to obtain a sorting result;
[0037] Intercepting a preset number of knowledge chunks from the sorting result to be determined as the retrieval result.
[0038] The second aspect of the present application discloses a retrieval enhancement processing system based on a large model, and the system includes:
[0039] An acquisition unit, configured to acquire a knowledge document corresponding to a user question when receiving the user question;
[0040] An analysis and segmentation unit, configured to perform adaptive analysis and segmentation on the knowledge document to obtain a segmentation result;
[0041] A recall unit, configured to perform retrieval recall on the segmentation result based on a multi-routing retrieval recall strategy and the user question to obtain a knowledge recall result; wherein, the multi-routing retrieval recall strategy is determined by sparse retrieval and dense retrieval;
[0042] A refined ranking unit, configured to refine the knowledge recall results based on a refined ranking model to obtain refined ranking results;
[0043] A sorting unit, configured to sort the refined ranking results according to a preset sorting method to determine the retrieval results.
[0044] Preferably, the analysis and segmentation unit includes:
[0045] An identification module, configured to identify each feature type in the knowledge document;
[0046] An analysis module, configured to call corresponding analysis methods according to the respective feature types to perform content analysis on the adaptive knowledge corpus of the knowledge document to obtain an analysis result;
[0047] A segmentation module, configured to perform corpus segmentation on the analysis result by combining a hierarchical structure and linear segmentation by character length to obtain a segmentation result.
[0048] Preferably, the segmentation module includes:
[0049] The first segmentation sub-module is used to segment the parsed result into each chapter content; wherein, the chapter content includes at least the chapter length and the chapter text content;
[0050] The determination sub-module is used to determine the lengths of all chapters from the chapter contents;
[0051] The first merging sub-module is used to merge the consecutive chapter contents corresponding to the consecutive same-level chapter lengths if the consecutive same-level chapter lengths among all chapter lengths are less than or equal to a preset threshold;
[0052] The second merging sub-module is used to merge the consecutive chapter contents of the same level if a single chapter length among all chapter lengths is less than or equal to the preset threshold;
[0053] The second segmentation sub-module is used to, if a single chapter length among all chapter lengths is greater than the preset threshold, judge the threshold length of the next-level directory of the current-level chapter until the single chapter length of the last-level directory is greater than the preset threshold, and complete the threshold segmentation according to the line break characters and punctuation marks of the chapter content, so as to complete the corpus segmentation of the parsed result and obtain the segmentation result.
[0054] Preferably, the recall unit includes:
[0055] The setting module is used to set the global cosine similarity threshold for knowledge search and the number of retrieved knowledge chunks;
[0056] The extraction module is used to extract the keywords of the user question;
[0057] The acquisition module is used to respectively perform vector representation on the user question and the keywords through an embedding model to obtain the sparse vector of the retrieval query, the dense vector of the retrieval query, the sparse vector of the retrieval keyword, and the dense vector of the retrieval keyword;
[0058] The first retrieval processing module is used to perform sparse vector retrieval on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge chunks corresponding to the first sparse vector, perform dense vector retrieval on the dense vector of the retrieval query to obtain the knowledge chunks corresponding to the first dense vector, and de-duplicate and merge the knowledge chunks corresponding to the first sparse vector and the knowledge chunks corresponding to the first dense vector to obtain the knowledge chunks matching the retrieval query;
[0059] The second retrieval processing module is used to perform sparse vector retrieval on the sparse vector of the retrieval keyword through an open-source vector database to obtain the knowledge block corresponding to the second sparse vector, perform dense vector retrieval on the dense vector of the retrieval keyword to obtain the knowledge block corresponding to the second dense vector, and de-duplicate and merge the knowledge block corresponding to the second sparse vector and the knowledge block corresponding to the second dense vector to obtain the knowledge block matching the retrieval keyword;
[0060] The third retrieval processing module is used to perform sparse vector retrieval for knowledge generation problem matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge block corresponding to the third sparse vector, perform dense vector retrieval for knowledge generation problem matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the third dense vector, and merge the knowledge block corresponding to the third sparse vector and the knowledge block corresponding to the third dense vector to obtain the knowledge block matching the knowledge generation problem;
[0061] The fourth retrieval processing module is used to perform sparse vector retrieval for abstract matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge block corresponding to the fourth sparse vector, perform dense vector retrieval for abstract matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the fourth dense vector, and de-duplicate and merge the knowledge block corresponding to the fourth sparse vector and the knowledge block corresponding to the fourth dense vector to obtain the knowledge block matching the abstract;
[0062] The calculation module is used to calculate the scores of the knowledge block matching the retrieval query, the knowledge block matching the retrieval keyword, the knowledge block matching the knowledge generation problem, and the knowledge block matching the abstract through the multi-route retrieval result formula to obtain the comprehensive score of all retrieval results;
[0063] The score filtering module is used to filter the comprehensive score of all retrieval results through the number of retrieved knowledge blocks and the global cosine similarity threshold of knowledge search to obtain the knowledge recall result.
[0064] As can be seen from the above technical solutions, the present application discloses a retrieval enhancement processing method and system based on a large model. When receiving a user question, the knowledge document corresponding to the user question is obtained, the knowledge document is adaptively parsed and segmented to obtain the segmentation result, and based on the multi-route retrieval recall strategy and the user question, the segmentation result is retrieved and recalled to obtain the knowledge recall result, wherein the multi-route retrieval recall strategy is determined by sparse retrieval and dense retrieval, the knowledge recall result is refined through a refinement model to obtain the refined result, and the refined result is sorted according to a preset sorting method to determine the retrieval result.
[0065] Through the above solution, for a large number of knowledge documents of different format types, batch adaptive parsing and segmentation can be achieved. For the segmented knowledge blocks in the segmentation results, a multi-routing retrieval and recall strategy based on hybrid retrieval is adopted. The multi-routing retrieval and recall strategy includes multiple recall dimensions such as knowledge matching, knowledge keyword matching, knowledge block generation problem matching, and document abstract matching. And through a multi-routing retrieval and recall strategy that combines sparse retrieval and dense retrieval, knowledge matching, knowledge keyword matching, knowledge generation-related problem matching, and knowledge abstract matching of user questions are realized. The refined ranking results obtained by using the refined ranking model trained with data in the communication field in cooperation with the multi-routing retrieval and recall strategy are sorted according to a preset sorting method to determine the retrieval results, that is, the user questions and knowledge blocks are scored, and the result with the highest score is returned. On the basis of having achieved knowledge recall, a secondary retrieval of questions and knowledge blocks is performed based on a semantic model to further improve the knowledge retrieval accuracy of the retrieval-enhanced generation Q&A system, thereby improving the accuracy of knowledge Q&A and intelligent dialogue content of the retrieval-enhanced generation Q&A system. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained according to the provided accompanying drawings.
[0067] Figure 1 It is a schematic flowchart of a retrieval enhancement processing method based on a large model disclosed in an embodiment of the present application;
[0068] Figure 2 It is a schematic diagram of the adaptive parsing and segmentation of knowledge documents disclosed in an embodiment of the present application;
[0069] Figure 3 It is an example diagram of the parsing of an OCR layout analysis model disclosed in an embodiment of the present application;
[0070] Figure 4 It is an example diagram of the segmentation method of text, chapters, titles, and tables of contents disclosed in an embodiment of the present application;
[0071] Figure 5 It is a flowchart of RAG knowledge Q&A disclosed in an embodiment of the present application;
[0072] Figure 6 It is a schematic structural diagram of a retrieval enhancement processing system based on a large model disclosed in an embodiment of the present application;
[0073] Figure 7 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. Detailed implementation manners
[0074] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0075] In the present application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the said element.
[0076] As can be seen from the background art, in the existing RAG question-and-answer system, since knowledge retrieval and recall usually use cosine similarity semantic retrieval based on a database, text content retrieval, etc., for situations where the user's question expression is relatively one-sided or the question itself belongs to knowledge details, etc., simple knowledge text matching cannot comprehensively and multi-facetedly recall the knowledge block where the knowledge point corresponding to the question is located, resulting in the failure of knowledge retrieval and recall of the RAG question-and-answer system, thereby reducing the accuracy of knowledge question-and-answer and intelligent dialogue content of the RAG question-and-answer system. Therefore, how to improve the knowledge retrieval of the RAG question-and-answer system to improve the accuracy of knowledge question-and-answer and intelligent dialogue content of the RAG question-and-answer system is an urgent problem to be solved in the present application.
[0077] To solve the above problems, this application discloses a retrieval enhancement processing method and system based on a large model. For a large number of knowledge documents in different format types, it can achieve batch adaptive parsing and segmentation. For the segmented knowledge chunks in the segmentation results, a multi-routing retrieval and recall strategy based on hybrid retrieval is adopted. The multi-routing retrieval and recall strategy includes multiple recall dimensions such as knowledge matching, knowledge keyword matching, knowledge chunk generation question matching, and document abstract matching. And through a multi-routing retrieval and recall strategy that combines sparse retrieval and dense retrieval to achieve knowledge matching, knowledge keyword matching, knowledge generation-related question matching, and knowledge abstract matching of user questions. The refined ranking results obtained by using a refined ranking model trained with communication field data in cooperation with the multi-routing retrieval and recall strategy are sorted according to a preset sorting method to determine the retrieval results, that is, score the user questions and knowledge chunks, and return the result with the highest score. On the basis of having achieved knowledge recall, a secondary retrieval of questions and knowledge chunks is performed based on a semantic model to further improve the knowledge retrieval accuracy of the retrieval enhancement generation question-answering system, thereby improving the accuracy of knowledge question-answering and intelligent dialogue content of the retrieval enhancement generation question-answering system. The specific implementation manner will be specifically described through the following embodiments.
[0078] The technical field of this application belongs to the category of artificial intelligence, mainly involving technical fields such as large models, machine learning, deep learning, and knowledge graphs, and can be used in scenarios such as robot intelligent dialogue and knowledge question-answering in the customer service field to improve the accuracy of knowledge question-answering and robot intelligent dialogue content.
[0079] Reference Figure 1 As shown, it is a retrieval enhancement processing method based on a large model disclosed in an embodiment of this application. The retrieval enhancement processing method based on a large model mainly includes the following steps:
[0080] S101: When receiving a user question, obtain the knowledge document corresponding to the user question.
[0081] The document formats of the knowledge documents corresponding to the user questions include but are not limited to Portable Document Format (PDF), word, table (excel), text (txt), HyperText Markup Language (html), csv, ppt, md, tex, jpg, png, etc.
[0082] S102: Perform adaptive parsing and segmentation on the knowledge document to obtain a segmentation result.
[0083] This solution supports the input of corpora in different document format types, and automatically identifies feature types such as whether there are chapters, titles, table of contents structures, pictures, paragraphs, tables, page numbers, and whether the document has a source document in the knowledge document.
[0084] The process of adaptively parsing and segmenting a knowledge document to obtain a segmentation result is shown in A1 - A3.
[0085] A1: Identify each feature type in the knowledge document.
[0086] Among them, the feature types include chapter titles, table of contents structures, pictures, paragraphs, tables, page numbers, and documents, etc.
[0087] A2: According to each feature type, call the corresponding parsing method to perform content parsing of the adaptive knowledge corpus on multiple to - be - processed documents of different format types, and obtain a parsing result.
[0088] In A2, based on different feature classifications, call corresponding parsing methods such as picture optical character recognition (OCR) parsing, table structure parsing, paragraph content parsing, chapter title and table of contents parsing, header and footer information parsing, etc., so as to cooperate with the corresponding format - type custom and open - source document loading components such as AamPDFLoader and AamWordLoader to complete the content parsing of the knowledge corpus. The types of document loading components are not specifically limited in this application.
[0089] A3: Perform corpus segmentation on the parsing result through a combination of hierarchical structure and character - length linear segmentation to obtain a segmentation result.
[0090] In A3, automatically adapt to corresponding custom and open - source text segmentation components such as CharpterTextSplitter, LineTextSplitter, DocSeqTextSplitter, etc. to complete the segmentation of corpus blocks. The types of text segmentation components are not specifically limited in this application.
[0091] The automatic recognition effect of this solution has reached the level of manual judgment to a certain extent and has been applied in multiple on - site scenarios.
[0092] To facilitate understanding of the process of adaptively parsing and segmenting a knowledge document to obtain a segmentation result, it is described in combination with Figure 2 and B1 - B5, Figure 2 Fig. shows a schematic diagram of the adaptive parsing and segmentation of a knowledge document.
[0093] In the prior art, if the document format of the knowledge document is of the PDF type, there is a PDF document parsing component AamPDFLoader, which embeds an OCR model and is used to process pdf files mainly composed of pictures. However, the general OCR model has poor effects in specific fields.
[0094] Therefore, in this solution, the fine-tuning training of the OCR layout analysis model is completed by using a large amount of on-site annotated knowledge corpus in the communication field (such as on-site management methods, operation guides, business process flowcharts, etc.), and it can realize the recognition of title-level directories, pictures, tables, paragraph texts, picture titles, table titles, header and footer information, etc. in PDF format documents in the communication field;
[0095] At the same time, corresponding parsing methods are used for the above-mentioned identified feature types respectively, and while extracting the text content, the semantic structure information of various features is retained as much as possible, providing an effective semantic guidance basis for the subsequent segmentation of knowledge corpus blocks.
[0096] For the parsing process of the OCR layout analysis model, please refer to Figure 3 as shown below.
[0097] Figure 3 In [reference], through the OCR layout analysis model, the recognition knowledge document is convolved to obtain various feature types, such as document titles, chapter titles, texts, images, tables, page numbers, etc.
[0098] B1: Perform text segmentation on the parsing results to obtain the content of each chapter; among them, the chapter content includes at least the chapter length and the chapter text content.
[0099] The text chapter title directory segmentation component CharpterTextSplitter adopts a combination of hierarchical structure and linear segmentation of character length to realize the text segmentation of knowledge text according to the chapter, title, directory structure, etc. of the first, second, and third levels. For the convenience of understanding its specific process, it is described in combination with Figure 4 as follows. Figure 4 shows an example diagram of the segmentation method of text, chapter, title and directory.
[0100] Figure 4 In [reference], according to the hierarchical structure of the knowledge document, the knowledge text is segmented and the title is segmented.
[0101] B2: Determine the length of each chapter from the content of each chapter.
[0102] B3: If the consecutive chapter lengths of the same level in each chapter length are less than or equal to the preset threshold (chunk_size), merge the consecutive chapter contents corresponding to the consecutive chapter lengths of the same level.
[0103] In B3, when the lengths of consecutive chapters at the same level are less than the preset threshold, the content of consecutive chapters is merged. When the length of a single chapter is greater than the preset threshold, the threshold length judgment is carried out for the next-level directory of the current-level chapter until the length of a single chapter in the last-level directory is still greater than the preset threshold. Then, according to the line breaks, punctuation marks, etc. in the chapter text content, the threshold segmentation is completed in sequence. Otherwise (that is, if the length of a single chapter among the lengths of each chapter is less than or equal to the preset threshold), the content of consecutive chapters at the same level is directly merged.
[0104] This segmentation and merging algorithm based on the hierarchical structure of chapter titles and the text character length not only ensures the content independence and semantic integrity of a single knowledge text block, but also retains the semantic connection between consecutive knowledge blocks under the same chapter. While ensuring a high matching effect for single knowledge blocks of detailed knowledge points, it improves the retrieval and matching ability of multiple knowledge blocks for summary questions involving the entire chapter and paragraph.
[0105] Among them, the preset threshold is set according to the actual situation, and this application does not make specific limitations.
[0106] B4: If the length of a single chapter among the lengths of each chapter is less than or equal to the preset threshold, the content of consecutive chapters at the same level is merged.
[0107] B5: If the length of a single chapter among the lengths of each chapter is greater than the preset threshold, the threshold length judgment is carried out for the next-level directory of the current-level chapter until the length of a single chapter in the last-level directory is greater than the preset threshold. Then, according to the line breaks and punctuation marks in the chapter content, the threshold segmentation is completed to complete the corpus segmentation of the parsing result and obtain the segmentation result.
[0108] While realizing the adaptive parsing and segmentation of the corpus, this solution also provides capabilities such as knowledge block keyword extraction, knowledge block related question generation, and document summary generation, providing more starting points and bases for subsequent multi-dimensional knowledge retrieval and recall, being able to improve the effect of retrieval and recall, and thus improving the question answering accuracy of the entire RAG question answering system.
[0109] The RAG question answering system mainly focuses on how to improve the accuracy of the knowledge question answering system based on large language models (LLMs). Currently, the research in this field mainly focuses on the following aspects:
[0110] Optimizing the retrieval algorithm: Knowledge block retrieval and recall based on the knowledge database is a key step in determining the cognitive enhancement effect. If the knowledge retrieval in the database fails to effectively recall valid information, even if the semantic understanding and answer generation capabilities of the large language model are very excellent, they will not be able to be exerted. Therefore, how to optimize the retrieval algorithm and improve the knowledge retrieval recall rate is an important research direction for improving the accuracy of the RAG knowledge question answering system.
[0111] Knowledge Base Management: Knowledge retrieval and recall based on the knowledge base determine the importance of knowledge base management. Once there are a large number of noisy corpus blocks in the knowledge base, it will greatly interfere with the efficiency and accuracy of knowledge retrieval and recall. By extracting and classifying the features of various types of knowledge in the knowledge base and performing retrieval and recall based on the features, it will help reduce the interference of noisy knowledge and improve the accuracy of retrieval results.
[0112] Text Segmentation: The size of the segmentation granularity of knowledge texts has an obvious impact on the matching of the embedding model. Generally, for knowledge documents with the same chapter and paragraph, different granularity segmentations will result in different matching effects of the embedding model. Especially for the retrieval of detailed questions, a finer granularity segmentation is more likely to obtain an accurate match, while for the retrieval of questions with summary and generalization, larger granularity segmentation blocks can more effectively explain the questions.
[0113] Knowledge Graph: A knowledge graph is a structured way of representing knowledge, which can provide rich background knowledge for the question-answering system. Currently, researchers mainly explore how to combine the knowledge graph with pre-trained models to improve the accuracy of the question-answering system.
[0114] Knowledge Text Keyword Extraction:
[0115] In the retrieval method, text keyword retrieval can effectively supplement the results of text retrieval and highlight the key information of the user's question. Considering system efficiency, this application compares the keyword extraction schemes of open-source small models and large models, and finally selects the more efficient open-source small model jionlp.keyphrase.ChineseKeyPhrasesExtractor. Through testing and evaluation on public datasets, the extraction accuracy is 2 percentage points lower than that of the large model Qwen14b, but the extraction efficiency is 20 times higher than that of Qwen14b. As one of the knowledge retrieval and recall routes, it can completely replace the large model for knowledge text keyword extraction. The open-source embedding model bge-m3 is used to represent the knowledge text keywords as vectors, and the open-source databases are ElasticSearch and Milvus. Among them, ElasticSearch is used to store the knowledge text text and the knowledge keyword text keywordtext; Milvus is used to store the knowledge vectors vector and the knowledge keyword vectors keywordvector.
[0116] Problems Related to Knowledge Text Generation:
[0117] In the retrieval stage, a series of possible questions are generated for each document. These questions are considered relevant to the document content and can represent the potential questions to be answered by the document. The user's questions are matched with these relevant questions of the knowledge text, so as to enhance the retrieval effect.
[0118] For system efficiency considerations, this application uses a small open-source model doc2query trained with data fine-tuning in the communication field, and compares its performance with that of the large model Qwen14B in generating relevant questions for knowledge text chunks. After fine-tuning, the open-source small model doc2query has a lower effect in evaluating the cosine similarity score between the extracted questions and user questions through on-site sample data, such as a decrease of up to 8 percentage points compared to the large model Qwen14B when the score is above 90 points. However, its extraction efficiency is more than 10 times higher than that of Qwen14B. As one of the knowledge retrieval recall routes, it can replace the large model to generate knowledge text-related questions. The relevant questions generated by knowledge chunks are stored separately in the knowledge chunk-related question table. The open-source embedding model bge-m3 is used to represent the knowledge-related questions as vectors. The database ElasticSearch is used to store the knowledge-related question text querytext, and Milvus is used to store the knowledge-related question vectors queryvector. At the same time, the knowledge chunk-related question table and the knowledge chunk table are associated and mapped based on the knowledge chunk chunk_id.
[0119] Knowledge document abstract generation:
[0120] In this application, a large model is used to generate full-document abstracts for knowledge documents. At the same time, the open-source embedding model bge-m3 is used to represent the document abstracts as vectors. The database ElasticSearch is used to store the document abstract text abstracttext, and Milvus is used to store the document abstract vectors abstractvector.
[0121] S103: Based on the multi-route retrieval recall strategy and the user's question, retrieve and recall the segmentation results to obtain the knowledge recall results; among them, the multi-route retrieval recall strategy is determined by sparse retrieval and dense retrieval.
[0122] Among them, the multi-route retrieval recall strategy is a multi-route knowledge retrieval recall algorithm that combines sparse retrieval and dense retrieval based on hybrid retrieval. The specific algorithm principle steps are as shown in (1) to (9) below:
[0123] (1) Set the default recall result quantity (similarity_topk) for the similarity matching of knowledge chunk vectors, the default recall result quantity (docs_topk) for the word segmentation matching of knowledge chunk text content, the global cosine similarity threshold (similarity) for knowledge search, and the number of knowledge chunks to be retrieved (topk).
[0124] (2) Extract keywords from the user's question, and use the embedding model bge-m3 to represent the user's question and the user's question keywords as vectors respectively, and output the query, query vector, query keyword, and query keyword vector. Among them, query keyword vector refers to the process of converting query keywords into vector representations.
[0125] The query vector includes the sparse vector and the dense vector of the retrieval query.
[0126] The query keyword vector includes the sparse vector and the dense vector of the retrieval keyword.
[0127] (3) Perform similarity search and text content search: Enter the sparse vector and the dense vector of the text into the Milvus database respectively, and perform sparse and dense vector retrieval on the user's question based on the Milvus database. The combination of the two vector retrieval results is specifically shown in formula (1):
[0128] (1)
[0129] Among them, is the vector score; is the overall adjustment factor, which is used to adjust the scale of the final score; is the weight of the sparse vector retrieval; is the weight of the dense vector retrieval; S is the score of the sparse vector retrieval result; D is the score of the dense vector retrieval result.
[0130] Arrange in descending order or ascending order, and obtain the similarity_topk recalled knowledge chunks and their matching scores; Use the text content word segmentation to match the inverted index of docs_search based on the distributed search and analysis engine (ElsticSearch) to search for the docs_topk recalled knowledge chunks and their matching scores.
[0131] (4) Filter the above two recall results according to the similarity threshold similarity, retain the knowledge chunks that meet the threshold requirements and return them.
[0132] To facilitate the understanding of the process of filtering the above two recall results according to the similarity threshold similarity, retaining the knowledge chunks that meet the threshold requirements and returning them, an example is given here to illustrate:
[0133] For example, 1. Recall similarity_topk. Knowledge block example: {'page_content': '2. Spring in Nature: The Awakening and Blooming of Life #2.1 Meteorological Changes: The Overture of Warmth and Vitality\nIn spring in nature, the temperature warms up and precipitation increases, playing the overture of the awakening of life.', 'page_keyword': 'Spring in Nature, Life, Awakening, Blooming, Meteorological Changes, Warmth, Vitality, Overture, Temperature Rise, Precipitation Increase, Playing','source': ' / work / tac-maas / file_loader / upload / corpus / xxxx / A Poetic Picture of the Recovery of All Things in the Rhythm of Spring.docx', 'data_id': '5dYP0I0BRmlw0i2ZYPam','vector': [0.026929408311843872, 0.03982673957943916, -0.06498610973358154, ……, 0.0076858592219650745, -0.01657874882221222],'similarity': 0.9125836};
[0134] 2. Recall docs_topk. Knowledge block example: {'page_content': '2. Spring in Nature: The Awakening and Blooming of Life #2.1 Meteorological Changes: The Overture of Warmth and Vitality #2.1.1 Temperature Rise\nAs the direct point of the sun moves northward, the temperature gradually warms up. The average daily temperature rises from below zero or a low temperature state in the cold winter to around 10℃ - 20℃. The morning mist no longer condenses into frost but turns into gentle water vapor and dissipates in the sun, and the air is filled with a fresh and humid atmosphere, indicating that the earth is about to wake up.\n2. Spring in Nature: The Awakening and Blooming of Life #2.1 Meteorological Changes: The Overture of Warmth and Vitality #2.1.2 Precipitation Increase\nPrecipitation in spring...', 'page_keyword': 'xxxxx','source': ' / work / tac-maas / file_loader / upload / corpus / xxxx / A Poetic Picture of the Recovery of All Things in the Rhythm of Spring.docx', 'data_id': 'zNYP0I0BRmlw0i2ZXvbw','score': 33.20973,'similarity': 0.9012284}.
[0135] 3. Filter according to the threshold: Filter out the knowledge blocks whose'similarity' value is greater than or equal to the similarity threshold similarity from each knowledge block recalled by the two paths and return them.
[0136] (5) For the multi-route problem, respectively perform text matching between the problem text and the knowledge text, keyword matching between the problem keywords and the knowledge keywords, generate relevant questions for the problem matching knowledge chunks, match the problem with the summary, etc., to complete the multi-route recall of steps (1) to (4).
[0137] Among them, the multi-route problem text matching the knowledge text means routing such as problem text matching knowledge text, problem keyword matching knowledge keyword, problem matching knowledge chunk generating relevant questions, and problem matching summary.
[0138] (6) For the relevant questions generated by the user problem matching the knowledge chunk, through the relevant question table of the knowledge chunk, the knowledge chunk table, and the knowledge chunk chunk_id, perform an association mapping to obtain the knowledge chunk corresponding to the recall of this route.
[0139] (7) For the result of the user problem matching the document summary recall, which is the knowledge document (to improve the retrieval recall efficiency, the number of knowledge documents recalled is generally set to top1), this solution needs to perform a re-retrieval and matching between the user problem and the knowledge within the scope of the knowledge documents recalled this time, as the recall result under this route.
[0140] (8) The formula for the multi-route retrieval result is shown in formula (2):
[0141] (2)
[0142] Among them, is the final score of the multi-route retrieval result; is the result score of the problem text matching the knowledge text; is the result score of the problem keyword matching the knowledge keyword; is the result score of the problem matching the knowledge chunk to generate relevant questions; is the result score of the problem matching the summary; and and and are the weights of the corresponding retrieval methods; is the overall adjustment factor, used to adjust the comprehensive score of all retrieval results; is the diversity adjustment factor, used to enhance or weaken the difference between different retrieval scores; is the diversity score of each retrieval score value, The expression of is shown in formula (3).
[0143] (3)
[0144] Among them, R is the number of retrieval routes; is the score value returned by each retrieval route.
[0145] (9) Dedup the above multi-route retrieval results by fusion_ranked to obtain the knowledge recall results.
[0146] Specifically, based on the multi-route retrieval recall strategy and the user's question, retrieve and recall the segmentation results to obtain the process of the knowledge recall results, as shown in C1-C7.
[0147] C1: Set the global cosine similarity threshold (similarity) for knowledge search and the number of retrieved knowledge chunks (topk).
[0148] C2: Extract the keywords of the user's question.
[0149] C3: Respectively perform vector representations on the user's question and the keywords through an embedding model to obtain the sparse vector of the retrieval query, the dense vector of the retrieval query, the sparse vector of the retrieval keyword, and the dense vector of the retrieval keyword.
[0150] The execution process and principle of C3 are the same as those of (2) above and can be referred to. Therefore, they will not be elaborated here.
[0151] C4: Perform sparse vector retrieval on the sparse vector of the retrieval query through an open-source vector database (Milvus) to obtain the knowledge chunks corresponding to the first sparse vector, perform dense vector retrieval on the dense vector of the retrieval query to obtain the knowledge chunks corresponding to the first dense vector, and de-duplicate and merge the knowledge chunks corresponding to the first sparse vector and the knowledge chunks corresponding to the first dense vector to obtain the knowledge chunks matching the retrieval query.
[0152] In C4, the sparse vector and the dense vector of the retrieval query are respectively input into the Milvus database, and sparse and dense vector retrievals are performed on the user's question based on the Milvus database to obtain the knowledge chunks corresponding to the first dense vector and the knowledge chunks matching the retrieval query.
[0153] For easy understanding, the process of fusing and ranking (fusion_ranked) and deduplicating the multi-route retrieval results is illustrated here with an example:
[0154] For example, take 2 types of recall results as an example as follows:
[0155] Data to be sorted:
[0156] Sorting 1: [a, b, c, d, e];
[0157] Sorting 2: [c, b, a, d, f];
[0158] Scoring process, default k = 60;
[0159] Sorting 1: a → 1 / 60, b → 1 / 61, c → 1 / 62, d → 1 / 63, e → 1 / 64;
[0160] Sorting 2: c → 1 / 60, b → 1 / 61, a → 1 / 62, d → 1 / 63, f → 1 / 64;
[0161] Aggregation:
[0162] a = 1 / 60 + 1 / 62 = 61 / 1860;
[0163] b = 1 / 61 + 1 / 61 = 2 / 61;
[0164] c = 1 / 60 + 1 / 62 = 61 / 1860;
[0165] d = 1 / 63 + 1 / 63 = 2 / 63;
[0166] e = 1 / 64 = 0.015625;
[0167] f = 1 / 64 = 0.015625;
[0168] That is,
[0169] a ≈ 0.0327956;
[0170] b ≈ 0.0327868;
[0171] c ≈ 0.0327956;
[0172] d ≈ 0.0317460;
[0173] e = 0.015625;
[0174] f = 0.015625;
[0175] Sorting result: a, c, b, d, e, f (when the scores are the same, the result sorted in Sorting 1 is preferred).
[0176] C5: Perform sparse vector retrieval on the sparse vector of the retrieval keyword through an open-source vector database to obtain the knowledge block corresponding to the second sparse vector, perform dense vector retrieval on the dense vector of the retrieval keyword to obtain the knowledge block corresponding to the second dense vector, and de-duplicate and merge the knowledge block corresponding to the second sparse vector and the knowledge block corresponding to the second dense vector to obtain the knowledge block matching the retrieval keyword.
[0177] C6: Perform sparse vector retrieval for knowledge generation problem matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge block corresponding to the third sparse vector, perform dense vector retrieval for knowledge generation problem matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the third dense vector, and merge the knowledge block corresponding to the third sparse vector and the knowledge block corresponding to the third dense vector to obtain the knowledge block for knowledge generation problem matching.
[0178] C7: Perform sparse vector retrieval for summary matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge block corresponding to the fourth sparse vector, perform dense vector retrieval for summary matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the fourth dense vector, and perform deduplication and merging on the knowledge block corresponding to the fourth sparse vector and the knowledge block corresponding to the fourth dense vector to obtain the knowledge block for summary matching.
[0179] The execution process and principle of C4 - C7 are the same as those of (4) to (6) above, which can be referred to and will not be elaborated here.
[0180] C8: Calculate the scores of the retrieval query matching knowledge block, retrieval keyword matching knowledge block, knowledge generation problem matching knowledge block, and summary matching knowledge block through the multi-route retrieval result formula to obtain the comprehensive score of all retrieval results.
[0181] The execution process of C8 can be referred to as shown in the above multi-route retrieval result formula (2) and will not be elaborated here.
[0182] C9: Filter the comprehensive scores of all retrieval results through the number of retrieved knowledge blocks and the global cosine similarity threshold of knowledge search to obtain the knowledge recall results.
[0183] In C9, score filtering refers to sorting in descending order or ascending order and taking out the knowledge blocks that are multiples of top_k of the number of retrieved knowledge blocks.
[0184] S104: Based on the re-rank model, perform re-rank on the knowledge recall results to obtain the re-rank results.
[0185] In S104, through the domain-customized re-rank (info-reranker) model fine-tuned with communication domain data, perform semantic relevance scoring on the knowledge blocks in the knowledge recall results and the user questions to obtain the respective semantic relevance scores, so as to complete the process of re-ranking the knowledge recall results.
[0186] Model re-ranking is performed in the last stage of the retrieval process, which is used to merge and sort the results from different retrieval systems to ensure that the documents most relevant to the user's question are ranked at the front. This solution uses the open-source info-reranker model fine-tuned with data in the communication field. This info-reranker model scores the semantic relevance between the knowledge chunks recalled by multi-route and the user's question, sorts them from high to low according to the scores, and intercepts the top_k knowledge chunks for result return. Specifically, as Figure 5 shown. Figure 5 It shows the RAG knowledge Q&A flowchart.
[0187] S105: Sort the re-ranking results according to a preset sorting method to determine the retrieval results.
[0188] In S105, sort each semantic relevance score in the re-ranking results according to a preset sorting method to obtain a sorting result, and intercept a preset number of knowledge chunks from the sorting result to determine the retrieval results.
[0189] The preset sorting method can be a high-to-low sorting method or a low-to-high sorting method. The specific form of the preset sorting method is not limited in this application.
[0190] This solution can achieve batch adaptive parsing and segmentation of a large number of knowledge documents of different format types on the customer site. It uses an OCR model retrained with data in the communication field to process pictures and picture-based pdf files. This method supports multi-route knowledge retrieval recall. Through a retrieval method that combines sparse retrieval and dense retrieval, it realizes knowledge matching of user questions, knowledge keyword matching, knowledge generation-related question matching, and knowledge summary matching. In addition, this method uses a re-ranking model trained with data in the communication field to score the user's question and knowledge chunks in combination with the multi-route retrieval recall results, and returns the results with higher scores. Based on the semantic model, it performs secondary retrieval on the questions and knowledge chunks on the basis of having achieved knowledge recall, further improving the accuracy of knowledge retrieval, thereby improving the overall effect of RAG Q&A.
[0191] This solution is based on the RAG retrieval enhancement technology of domain re-ranking models and multi-route retrieval recall. It can achieve batch adaptive parsing and segmentation of a large number of knowledge documents of different format types. It uses an OCR model retrained with data in the communication field to process pictures and picture-based pdf files. For the segmented knowledge chunks, it adopts a multi-route retrieval recall strategy based on hybrid retrieval, including multiple recall dimensions such as knowledge matching, knowledge keyword matching, knowledge chunk generation question matching, and document summary matching. Finally, it uses the re-ranking model info-reranker fine-tuned with data in the communication field to perform secondary re-ranking on the above multi-route knowledge recall results, sorts them in descending order according to the model re-ranking scores, and takes the top N results with higher scores as the final retrieval results.
[0192] This solution solves the following problems:
[0193] 1. Adaptive document parsing and segmentation: For knowledge documents of different format types, automatically identify document features such as chapter titles, pictures, tables, etc., and execute corresponding parsing and segmentation strategies to achieve automatic parsing and segmentation of various types of knowledge corpora.
[0194] 2. Optimization of text parsing in specific domains: For documents in specific formats, such as pictures and PDF files, use domain-specific OCR models to improve the text parsing effect, especially in professional fields such as the communication field, to improve the effect of knowledge data storage and user experience.
[0195] 3. Multi-route retrieval and recall strategy: Build a multi-route knowledge retrieval and recall strategy based on hybrid retrieval, combine text content matching and semantic matching, adopt a combination of sparse retrieval and hybrid retrieval, and improve the accuracy and comprehensiveness of retrieval and recall through knowledge retrieval and recall in multiple dimensions.
[0196] 4. Multi-dimensional knowledge retrieval and recall: Incorporate multiple dimensions such as knowledge matching, keyword matching, question matching, and document abstract matching into the retrieval and recall routes to ensure the integrity of the recall results.
[0197] 5. Model fine-tuning strategy: Based on excellent long-text and multilingual fine-tuning models in the industry, develop an open-source info-reranker model for the communication field, and perform secondary scoring and evaluation on the recall results to improve the accuracy of the final knowledge retrieval.
[0198] Key technical value points of this solution:
[0199] 1. Adaptive document parsing: This solution uses an innovative adaptive parsing strategy that can intelligently identify and process various document formats, including but not limited to PDF, Word, Excel, and HTML documents, enhancing the flexibility and comprehensiveness of document parsing.
[0200] 2. Multi-feature recognition technology: By automatically identifying features such as chapter titles, pictures, paragraphs, tables, page numbers, and document sources in documents, the patent solution improves the accuracy and depth of document processing.
[0201] 3. Communication field customized OCR model: Based on data in the communication field, train an OCR model. Compared with traditional general OCR models, this model shows higher parsing efficiency and accuracy when processing images and documents in the communication field.
[0202] 4. Customized parsing tools: Components such as AamPDFLoader, AamWordLoader, AamExcelLoader, and AamHTMLLoader proposed in this application provide customized parsing tools for different types of documents, ensuring the efficiency and accuracy of the parsing process.
[0203] 5. Advanced text splitting components: The application of text splitting components such as CharpterTextSplitter, LineTextSplitter, and DocSeqTextSplitte ensures the semantic integrity of the text blocks after splitting and the logical coherence within the chapters.
[0204] 6. Hybrid retrieval algorithm: The patent implements a knowledge retrieval method based on the hybrid hybrid retrieval algorithm, which improves the effective recall rate of knowledge blocks related to the problem through multi-dimensional matching.
[0205] 7. Specific domain problem rewriting model: The doc2query model is trained based on data in the communication domain. This model is specifically designed to generate relevant questions from knowledge blocks, optimizing the knowledge generation problem matching process and improving the application value and efficiency in the communication domain.
[0206] 8. Domain customized fine-tuning model: The info-reranker model fine-tuned based on data in the communication domain is used for secondary fine-tuning, further improving the accuracy and relevance of knowledge retrieval.
[0207] 9. Significant improvement in recall rate: Through the optimized retrieval method, a significant improvement in the recall rate is achieved, especially in terms of the recall rates at top1, top5, and top10, with quantitative improvements.
[0208] The above-mentioned value points reflect the important contributions of this solution in terms of technological innovation, efficiency improvement, accuracy enhancement, and user experience optimization, and have obvious market application potential and competitive advantages.
[0209] Key protection points of this solution:
[0210] 1. Communication domain customized OCR solution: This solution proposes an OCR model specifically designed for the communication domain. Through training on communication domain datasets, the layout analysis and image parsing processes are optimized. Compared with traditional general OCR models, this model shows higher parsing efficiency and accuracy when processing images and documents in the communication domain.
[0211] 2. Multi-dimensional Retrieval Technology: The patent includes a hybrid retrieval technology that combines sparse vectors, dense vectors, and text retrieval methods. This technology retrieves problems from multiple dimensions, including knowledge text matching, keyword matching, question generation matching, and document abstract matching, achieving parallel processing of retrieval tasks and enhancing the comprehensiveness and efficiency of retrieval.
[0212] 3. Specific Domain Question Rewriting Model: The patent includes a doc2query model trained on communication domain data. This model is specifically designed to generate relevant questions from knowledge chunks, optimizing the knowledge generation question matching process and improving the application value and efficiency within the communication domain.
[0213] 4. Specific Domain Fine-tuning Model: The patent covers an info-reranker model fine-tuned on communication domain data. This model performs in-depth secondary fine-tuning on the recall results, significantly enhancing the relevance and accuracy of retrieval results and optimizing the user retrieval experience.
[0214] It should be noted that various models, components, tools, etc. of this solution are all open-source, allowing users to freely use, copy, modify, and distribute software, etc.
[0215] The beneficial effects of the embodiments of this application: For a large number of knowledge documents in different format types, batch adaptive parsing and segmentation can be achieved. For the segmented knowledge chunks in the segmentation results, a multi-routing retrieval recall strategy based on hybrid retrieval is adopted. The multi-routing retrieval recall strategy includes multiple recall dimensions such as knowledge matching, knowledge keyword matching, knowledge chunk generation question matching, and document abstract matching, and the knowledge matching, knowledge keyword matching, knowledge generation-related question matching, and knowledge abstract matching of user questions are realized through a multi-routing retrieval recall strategy that combines sparse retrieval and dense retrieval. The fine-tuning model trained with communication domain data is used to cooperate with the multi-routing retrieval recall strategy to obtain the fine-tuning results, and the fine-tuning results are sorted according to the preset sorting method to determine the retrieval results, that is, to score the user questions and knowledge chunks and return the result with the highest score. On the basis of realizing knowledge recall, secondary retrieval of questions and knowledge chunks is performed based on the semantic model, further improving the knowledge retrieval accuracy of the retrieval-enhanced generation question-answering system, thereby improving the accuracy of knowledge question-answering and intelligent dialogue content of the retrieval-enhanced generation question-answering system.
[0216] Based on the above embodiments Figure 1 A retrieval-enhanced processing method based on a large model disclosed, the embodiments of this application also correspondingly disclose a retrieval-enhanced processing system based on a large model, as Figure 6 shown, the retrieval-enhanced processing system based on a large model includes:
[0217] An acquisition unit 601, configured to acquire a knowledge document corresponding to a user question when receiving the user question;
[0218] A parsing and splitting unit 602, configured to perform adaptive parsing and splitting on the knowledge document to obtain a splitting result;
[0219] A recall unit 603, configured to perform retrieval and recall on the splitting result based on a multi-routing retrieval and recall strategy and the user question to obtain a knowledge recall result; wherein, the multi-routing retrieval and recall strategy is determined by sparse retrieval and dense retrieval;
[0220] A fine-ranking unit 604, configured to perform fine-ranking on the knowledge recall result based on a fine-ranking model to obtain a fine-ranking result;
[0221] A sorting unit 605, configured to sort the fine-ranking result according to a preset sorting method to determine a retrieval result.
[0222] Further, the parsing and splitting unit 602 includes:
[0223] An identification module, configured to identify each feature type in the knowledge document;
[0224] A parsing module, configured to call corresponding parsing methods according to each feature type to perform content parsing of the adaptive knowledge corpus on the knowledge document to obtain a parsing result;
[0225] A splitting module, configured to perform corpus splitting on the parsing result by combining hierarchical structure and linear splitting by character length to obtain a splitting result.
[0226] Further, the splitting module includes:
[0227] A first sub-splitting module, configured to perform text splitting on the parsing result to obtain the content of each chapter; wherein, the chapter content includes at least the chapter length and the chapter text content;
[0228] A determination sub-module, configured to determine the length of each chapter from the content of each chapter;
[0229] A first merging sub-module, configured to merge the continuous chapter content corresponding to the continuous same-level chapter lengths if the continuous same-level chapter lengths in each chapter length are less than or equal to a preset threshold;
[0230] A second merging sub-module, configured to merge the continuous chapter content at the same level if the length of a single chapter in each chapter length is less than or equal to a preset threshold;
[0231] The second slicing sub-module is used to, if the length of a single chapter among the lengths of each chapter is greater than a preset threshold, perform threshold length judgment on the next-level directory of the current-level chapter until the length of a single chapter in the last-level directory is greater than the preset threshold, and complete threshold slicing according to the line break characters and punctuation marks in the chapter content, so as to perform corpus slicing on the parsing result to obtain a slicing result.
[0232] Further, the recall unit 603 includes:
[0233] A setting module for setting the global cosine similarity threshold of knowledge search and the number of retrieved knowledge chunks;
[0234] An extraction module for extracting keywords of the user's question;
[0235] An acquisition module for respectively performing vector representation on the user's question and keywords through an embedding model to obtain a sparse vector of the retrieval query, a dense vector of the retrieval query, a sparse vector of the retrieval keyword, and a dense vector of the retrieval keyword;
[0236] The first retrieval processing module is used to perform sparse vector retrieval on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge chunks corresponding to the first sparse vector, perform dense vector retrieval on the dense vector of the retrieval query to obtain the knowledge chunks corresponding to the first dense vector, and perform deduplication and merging on the knowledge chunks corresponding to the first sparse vector and the knowledge chunks corresponding to the first dense vector to obtain the knowledge chunks matching the retrieval query;
[0237] The second retrieval processing module is used to perform sparse vector retrieval on the sparse vector of the retrieval keyword through an open-source vector database to obtain the knowledge chunks corresponding to the second sparse vector, perform dense vector retrieval on the dense vector of the retrieval keyword to obtain the knowledge chunks corresponding to the second dense vector, and perform deduplication and merging on the knowledge chunks corresponding to the second sparse vector and the knowledge chunks corresponding to the second dense vector to obtain the knowledge chunks matching the retrieval keyword;
[0238] The third retrieval processing module is used to perform sparse vector retrieval for knowledge generation question matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge chunks corresponding to the third sparse vector, perform dense vector retrieval for knowledge generation question matching on the dense vector of the retrieval query to obtain the knowledge chunks corresponding to the third dense vector, and merge the knowledge chunks corresponding to the third sparse vector and the knowledge chunks corresponding to the third dense vector to obtain the knowledge chunks matching the knowledge generation question;
[0239] The fourth retrieval processing module is used to perform sparse vector retrieval for summary matching of the sparse vector of the retrieval query through an open-source vector database, obtain the knowledge block corresponding to the fourth sparse vector, perform dense vector retrieval for summary matching of the dense vector of the retrieval query, obtain the knowledge block corresponding to the fourth dense vector, and deduplicate and merge the knowledge block corresponding to the fourth sparse vector and the knowledge block corresponding to the fourth dense vector to obtain the summary matching knowledge block;
[0240] The calculation module is used to calculate the scores of the retrieval query matching knowledge block, the retrieval keyword matching knowledge block, the knowledge generation question matching knowledge block, and the summary matching knowledge block through the multi-routing retrieval result formula to obtain the comprehensive score of all retrieval results;
[0241] The score filtering module is used to filter the comprehensive scores of all retrieval results through the number of retrieved knowledge blocks and the global cosine similarity threshold of knowledge search to obtain the knowledge recall result.
[0242] Furthermore, the fine-ranking unit 604 is specifically used to score the semantic relevance between the knowledge blocks in the knowledge recall result and the user question through a fine-ranking model fine-tuned with communication field data, obtain each semantic relevance score, so as to complete the process of fine-ranking the knowledge recall result based on the fine-ranking model to obtain the fine-ranking result.
[0243] Furthermore, the sorting unit 605 includes:
[0244] The sorting module is used to sort each semantic relevance score in the fine-ranking result according to a preset sorting method to obtain the sorting result;
[0245] The determination module is used to intercept a preset number of knowledge blocks from the sorting result and determine them as the retrieval result.
[0246] Advantages of the embodiments of the present application: For a large number of knowledge documents of different format types, batch adaptive parsing and segmentation can be achieved. For the segmented knowledge chunks in the segmentation results, a multi-routing retrieval and recall strategy based on hybrid retrieval is adopted. The multi-routing retrieval and recall strategy includes multiple recall dimensions such as knowledge matching, knowledge keyword matching, knowledge chunk generation problem matching, and document abstract matching. And through a multi-routing retrieval and recall strategy that combines sparse retrieval and dense retrieval, knowledge matching, knowledge keyword matching, knowledge generation-related problem matching, and knowledge abstract matching of user questions are realized. The refined ranking results obtained by the refined ranking model trained with data in the communication field in cooperation with the multi-routing retrieval and recall strategy are sorted according to a preset sorting method to determine the retrieval results, that is, the user questions and knowledge chunks are scored, and the result with the highest score is returned. On the basis of having achieved knowledge recall, secondary retrieval of questions and knowledge chunks is performed based on a semantic model to further improve the knowledge retrieval accuracy of the retrieval-enhanced generation Q&A system, thereby improving the accuracy of knowledge Q&A and intelligent dialogue content of the retrieval-enhanced generation Q&A system.
[0247] The embodiments of the present application also provide a storage medium, where the storage medium includes stored instructions, and when the instructions run, the device where the storage medium is located is controlled to execute the retrieval-enhanced processing method based on a large model as described above.
[0248] The embodiments of the present application also provide an electronic device, and its structural schematic diagram is as Figure 7 shown, specifically including a memory 701 and one or more instructions 702, where one or more instructions 702 are stored in the memory 701 and are configured to be executed by one or more processors 703 to execute the one or more instructions 702 to execute the retrieval-enhanced processing method based on a large model as described above.
[0249] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0250] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0251] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0252] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0253] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0254] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A retrieval enhancement processing method based on a large model, characterized in that, The method includes: When receiving a user question, obtaining the knowledge document corresponding to the user question; Performing adaptive parsing and segmentation on the knowledge document to obtain a segmentation result; Based on a multi-route retrieval and recall strategy and the user question, retrieving and recalling the segmentation result to obtain a knowledge recall result; wherein, the multi-route retrieval and recall strategy is determined by sparse retrieval and dense retrieval; Performing fine ranking on the knowledge recall result based on a fine ranking model to obtain a fine ranking result; Sorting the fine ranking result according to a preset sorting method to determine a retrieval result.
2. The method according to claim 1, wherein The performing adaptive parsing and segmentation on the knowledge document to obtain a segmentation result includes: Identifying each feature type in the knowledge document; According to each feature type, calling a corresponding parsing method to perform content parsing on the adaptive knowledge corpus of the knowledge document to obtain a parsing result; Performing corpus segmentation on the parsing result by combining a hierarchical structure and linear segmentation by character length to obtain a segmentation result.
3. The method according to claim 2, wherein The performing corpus segmentation on the parsing result by combining a hierarchical structure and linear segmentation by character length to obtain a segmentation result includes: Performing text segmentation on the parsing result to obtain the content of each chapter; wherein, the content of each chapter includes at least the chapter length and the chapter text content; Determining the length of each chapter from the content of each chapter; If the consecutive chapter lengths at the same level in the lengths of each chapter are less than or equal to a preset threshold, merging the consecutive chapter contents corresponding to the consecutive chapter lengths at the same level; If the length of a single chapter in the lengths of each chapter is less than or equal to the preset threshold, merging the consecutive chapter contents at the same level; If the length of a single chapter in the lengths of each chapter is greater than the preset threshold, performing a threshold length judgment on the next-level directory of the current-level chapter until the length of a single chapter in the last-level directory is greater than the preset threshold, and completing threshold segmentation according to the line break characters and punctuation marks in the chapter content to complete the corpus segmentation of the parsing result and obtain a segmentation result.
4. The method according to claim 1, characterized in that, The retrieving and recalling the segmentation result based on a multi-route retrieval and recall strategy and the user question to obtain a knowledge recall result includes: Setting a global cosine similarity threshold for knowledge search and the number of retrieved knowledge chunks; Extracting the keywords of the user question; Performing vector representation on the user question and the keywords respectively through an embedding model to obtain a sparse vector of the retrieval query, a dense vector of the retrieval query, a sparse vector of the retrieval keyword, and a dense vector of the retrieval keyword; Performing sparse vector retrieval on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge chunks corresponding to the first sparse vector, performing dense vector retrieval on the dense vector of the retrieval query to obtain the knowledge chunks corresponding to the first dense vector, and de-duplicating and merging the knowledge chunks corresponding to the first sparse vector and the knowledge chunks corresponding to the first dense vector to obtain the knowledge chunks matching the retrieval query; Perform sparse vector retrieval on the sparse vector of the retrieval keyword through an open-source vector database to obtain the knowledge block corresponding to the second sparse vector, perform dense vector retrieval on the dense vector of the retrieval keyword to obtain the knowledge block corresponding to the second dense vector, and deduplicate and merge the knowledge block corresponding to the second sparse vector and the knowledge block corresponding to the second dense vector to obtain the knowledge block matching the retrieval keyword; Perform sparse vector retrieval for knowledge generation problem matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge block corresponding to the third sparse vector, perform dense vector retrieval for knowledge generation problem matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the third dense vector, and merge the knowledge block corresponding to the third sparse vector and the knowledge block corresponding to the third dense vector to obtain the knowledge block matching the knowledge generation problem; Perform sparse vector retrieval for abstract matching on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge block corresponding to the fourth sparse vector, perform dense vector retrieval for abstract matching on the dense vector of the retrieval query to obtain the knowledge block corresponding to the fourth dense vector, and deduplicate and merge the knowledge block corresponding to the fourth sparse vector and the knowledge block corresponding to the fourth dense vector to obtain the knowledge block matching the abstract; Calculate the scores of the knowledge blocks matching the retrieval query, the knowledge blocks matching the retrieval keyword, the knowledge blocks matching the knowledge generation problem, and the knowledge blocks matching the abstract through a multi-route retrieval result formula to obtain the comprehensive score of all retrieval results; Perform score filtering on the comprehensive score of all retrieval results through the number of retrieved knowledge blocks and the global cosine similarity threshold of knowledge search to obtain the knowledge recall result.
5. The method according to claim 1, wherein Performing fine ranking on the knowledge recall result based on the fine-ranking model to obtain the fine-ranking result, including: Using the fine-ranking model fine-tuned with communication domain data to score the semantic relevance between the knowledge blocks in the knowledge recall result and the user question, obtaining each semantic relevance score, so as to complete the process of performing fine ranking on the knowledge recall result based on the fine-ranking model to obtain the fine-ranking result.
6. The method according to claim 1, wherein Sorting the fine-ranking result according to a preset sorting method to determine the retrieval result, including: Sorting the semantic relevance scores in the fine-ranking result according to a preset sorting method to obtain the sorting result; Intercepting a preset number of knowledge blocks from the sorting result to determine the retrieval result.
7. A retrieval enhancement processing system based on a large model, characterized in that, The system includes: An acquisition unit for acquiring the knowledge document corresponding to the user question when receiving the user question; An analysis and segmentation unit for adaptively analyzing and segmenting the knowledge document to obtain the segmentation result; A recall unit for retrieving and recalling the segmentation result based on a multi-route retrieval recall strategy and the user question to obtain the knowledge recall result; wherein, the multi-route retrieval recall strategy is determined by sparse retrieval and dense retrieval; A fine-ranking unit for performing fine ranking on the knowledge recall result based on the fine-ranking model to obtain the fine-ranking result; A sorting unit for sorting the fine-ranking result according to a preset sorting method to determine the retrieval result.
8. The system according to claim 7, wherein The analysis and segmentation unit includes; An identification module for identifying each feature type in the knowledge document; An analysis module for adaptively analyzing the content of the knowledge corpus of the knowledge document by calling corresponding analysis methods according to the respective feature types to obtain an analysis result; A segmentation module for segmenting the analysis result by combining hierarchical structure and linear segmentation by character length to obtain a segmentation result.
9. The system according to claim 8, wherein The segmentation module includes: A first sub-segmentation module for text-segmenting the analysis result to obtain the content of each chapter; wherein, the chapter content includes at least the chapter length and the chapter text content; A determination sub-module for determining the length of each chapter from the content of each chapter; A first merging sub-module for merging the consecutive chapter contents corresponding to the consecutive same-level chapter lengths if the consecutive same-level chapter lengths in the lengths of each chapter are less than or equal to a preset threshold; A second merging sub-module for merging the consecutive chapter contents at the same level if the length of a single chapter in the lengths of each chapter is less than or equal to the preset threshold; A second sub-segmentation module for, if the length of a single chapter in the lengths of each chapter is greater than the preset threshold, determining the threshold length for the next-level directory of the current-level chapter until the length of a single chapter in the last-level directory is greater than the preset threshold, and performing threshold segmentation according to the line break characters and punctuation marks in the chapter content to complete the corpus segmentation of the analysis result and obtain a segmentation result.
10. The system according to claim 7, characterized in that, The recall unit includes: A setting module for setting the global cosine similarity threshold for knowledge search and the number of retrieved knowledge chunks; An extraction module for extracting the keywords of the user question; An acquisition module for respectively performing vector representation on the user question and the keywords through an embedding model to obtain a sparse vector of the retrieval query, a dense vector of the retrieval query, a sparse vector of the retrieval keyword, and a dense vector of the retrieval keyword; A first retrieval processing module for performing sparse vector retrieval on the sparse vector of the retrieval query through an open-source vector database to obtain the knowledge chunks corresponding to the first sparse vector, performing dense vector retrieval on the dense vector of the retrieval query to obtain the knowledge chunks corresponding to the first dense vector, and de-duplicating and merging the knowledge chunks corresponding to the first sparse vector and the knowledge chunks corresponding to the first dense vector to obtain the knowledge chunks matching the retrieval query; A second retrieval processing module for performing sparse vector retrieval on the sparse vector of the retrieval keyword through an open-source vector database to obtain the knowledge chunks corresponding to the second sparse vector, performing dense vector retrieval on the dense vector of the retrieval keyword to obtain the knowledge chunks corresponding to the second dense vector, and de-duplicating and merging the knowledge chunks corresponding to the second sparse vector and the knowledge chunks corresponding to the second dense vector to obtain the knowledge chunks matching the retrieval keyword; The third retrieval processing module is used to perform sparse vector retrieval for knowledge generation problem matching on the sparse vector of the retrieval query through an open-source vector database, obtain the knowledge block corresponding to the third sparse vector, perform dense vector retrieval for knowledge generation problem matching on the dense vector of the retrieval query, obtain the knowledge block corresponding to the third dense vector, and merge the knowledge block corresponding to the third sparse vector and the knowledge block corresponding to the third dense vector to obtain the knowledge block for knowledge generation problem matching; The fourth retrieval processing module is used to perform sparse vector retrieval for abstract matching on the sparse vector of the retrieval query through an open-source vector database, obtain the knowledge block corresponding to the fourth sparse vector, perform dense vector retrieval for abstract matching on the dense vector of the retrieval query, obtain the knowledge block corresponding to the fourth dense vector, and perform deduplication and merging on the knowledge block corresponding to the fourth sparse vector and the knowledge block corresponding to the fourth dense vector to obtain the knowledge block for abstract matching; The calculation module is used to calculate the scores of the retrieval query matching knowledge block, the retrieval keyword matching knowledge block, the knowledge generation problem matching knowledge block, and the abstract matching knowledge block through a multi-route retrieval result formula to obtain the comprehensive score of all retrieval results; The score filtering module is used to perform score filtering on the comprehensive score of all retrieval results through the number of retrieved knowledge blocks and the global cosine similarity threshold of knowledge search to obtain the knowledge recall result.
Citation Information
Patent Citations
LLM-based intelligent questioning and answering method, system and equipment in power field and medium
CN118277521A
RAG-based vertical domain knowledge multi-round question and answer method
CN118964556A
Multipath fusion-based law and regulation recommendation system and method
CN119903234A
A system for generating answers to multiple questions using rag-based generative artificial intelligence technology
KR102710159B1
Cited By
Retrieval enhancement generation method and device
CN120541206A
Multi-modal mixed retrieval enhanced question and answer method and device
CN121029927A
Long document understanding method and device, electronic equipment and storage medium
CN121170820A
Intelligent questioning and answering system for road traffic safety laws and regulations based on retrieval enhancement generation
CN122285855A