Knowledge base generation and retrieval method based on RAG technology and intelligent question answering system
By constructing efficient vocabulary indexes in knowledge base generation and search methods and combining large language model optimization generation results, the efficiency problem of vector search in large-scale data sets is solved, and high-quality answer generation and information retrieval is achieved.
Patent Information
- Application Number
- CN202411927050.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
Existing knowledge base generation and search methods based on RAG technology. In large-scale data sets and high-dimensional spaces, it is difficult for vector search to efficiently retrieve the most relevant information, affecting the quality of generated answers.
By collecting documents from diverse data sources in multiple professional fields and performing format conversion, preprocessing, tiling and vectorization, we can build efficient vocabulary indexes, including inverted indexes and semantic indexes, and achieve hierarchical retrieval and keyword matching. Combining the optimization generation results of the large language model (LLM) optimization, the reasoning ability of the large model is further improved through chain reasoning and prompt word optimization modules.
It improves the accuracy and real-time nature of information processing, ensures that the generated answers are highly accurate and logical, and is suitable for fields such as intelligent question and answer, knowledge base construction and real-time information retrieval.
Smart Images

Figure CN120045653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a knowledge base generation and retrieval method and an intelligent question-answering system based on RAG technology. Background Art
[0002] Large language models have shown great potential in text generation and semantic understanding, but they rely on a fixed pre-trained knowledge base, which leads to timeliness and accuracy limitations when dealing with rapidly changing and domain-specific information. To solve this problem, RAG (Retrieval-Augmented Generation) technology combines information retrieval and generation models, and can dynamically retrieve external data when generating answers, thereby making up for the shortcomings of language models and improving the accuracy and comprehensiveness of answers.
[0003] In practical applications, relying solely on a single vector search often cannot meet diverse needs. Nevertheless, RAG methods based on vector retrieval still face challenges, especially in large-scale datasets and high-dimensional spaces, where vector search may not be able to efficiently retrieve the most relevant information, affecting the quality of generated answers. Summary of the invention
[0004] To solve the above problems, the present invention provides a knowledge base generation and retrieval method and an intelligent question-answering system based on RAG technology. The present invention is implemented through the following technical solutions.
[0005] A knowledge base generation method based on RAG technology comprises the following steps:
[0006] S101, collects documents in different formats from diverse data sources in multiple professional fields;
[0007] S102, document format conversion and preprocessing:
[0008] Format conversion: convert documents in different formats (PDF, Word, HTML, etc.) into text data;
[0009] Text preprocessing: semantic error correction, word segmentation, stop word removal, and noise data cleaning of text data;
[0010] S103, slicing the preprocessed text data into multiple logical segments, evaluating the response quality according to different slicing sizes, finding the best slicing granularity, and using vectorization technology to convert the text content into a high-dimensional vector;
[0011] S104, generating hypothetical questions, generating hypothetical questions based on the segmented text to simulate the user's query scenario;
[0012] S105, Lexical Index Construction, constructing an efficient lexical index, including the following operations:
[0013] Inverted index, associating keywords with the positions in the document content to achieve rapid positioning;
[0014] Semantic index, combined with vector representation, supporting semantic similarity retrieval and enhancing the relevance of retrieval results;
[0015] S106, Layering the text library, generating abstracts for documents using text summarization techniques; generating vectors and keywords for the full text content of the text;
[0016] S107, Evaluating the quality of retrieval results and dynamically adjusting the processing flow.
[0017] Furthermore, in the step S107, it includes:
[0018] Quality evaluation, using automated evaluation metrics and manual review to evaluate the relevance and accuracy of retrieval results;
[0019] Dynamic adjustment, according to the evaluation results, optimizing the retrieval algorithm, vectorization method or text chunking strategy to ensure that subsequent steps can provide higher-quality results;
[0020] Continuous optimization, continuously providing feedback and adjusting the text processing flow.
[0021] A retrieval method based on the RAG technology, including the following specific methods:
[0022] S201, Vector Retrieval, converting the user query semantics into a high-dimensional vector representation and comparing it with the high-dimensional semantic vectors converted from text data. Through similarity calculation, matching the document set most relevant to the user query semantics;
[0023] S202, Keyword Retrieval, matching the keywords in the user query and quickly locating the documents containing these words, screening the document set relevant to the user query; for questions with weak semantic extensibility, further retrieve in combination with vectors;
[0024] S203, Hierarchical Retrieval, preliminarily screening the abstract part of the document to quickly narrow down the range of candidate documents; in the documents screened by the abstract, conducting in-depth analysis on the full text to mine core information and detailed content;
[0025] S204, Question-Question Retrieval, parsing the context and keywords of the user question. Combining the query context, converting the question into a structured retrieval statement; through the semantic matching model, locating the documents or answers relevant to the user question;
[0026] S205, Rearranging after Retrieval, applying an advanced sorting model to perform secondary sorting on the preliminary results;
[0027] S206, verifying the search results, using an external verification system to perform consistency check on the search results;
[0028] S207, deduplicate the information, automatically identify duplicate content through semantic similarity detection, remove redundant information, and reorganize the output according to the importance of the content.
[0029] Furthermore, in step S202, the keyword search includes keyword extraction optimization, multi-keyword combination matching and keyword weight adjustment.
[0030] Furthermore, in S205, rearrangement after retrieval includes:
[0031] Preliminary sorting: Use traditional sorting algorithms to preliminarily sort candidate documents and generate a basic ranking list based on the matching degree between the query and the document;
[0032] Advanced sorting: Use advanced sorting models to optimize the preliminary sorting results and adjust the document priority based on indicators such as semantic relevance, document authority, and content completeness;
[0033] Semantic relevance optimization: During the sorting process, the semantic relevance between the document and the query is re-evaluated in combination with the contextual information of the user's query, and the most relevant and contextual document content is displayed first;
[0034] Personalized sorting: personalize the sorting results based on the user's historical query records and preferences.
[0035] Furthermore, in step S206, verifying the search results includes:
[0036] Use external knowledge base or verification model to check the semantic consistency of retrieval results.
[0037] Retrieve content through manual review and automated indicator evaluation to eliminate irrelevant or noisy data.
[0038] An intelligent question-answering system based on RAG technology includes the following modules:
[0039] Document preprocessing module, which is used to collect documents from data sources in multiple professional fields and perform formatting and semantic cleaning;
[0040] Retrieval module, including vector retrieval, keyword retrieval and hierarchical retrieval functions, used to quickly locate relevant documents;
[0041] The question analysis module analyzes the questions input by users, converts them into high-dimensional semantic vectors and generates structured query statements;
[0042] Prompt word optimization module, used to optimize the generation process;
[0043] A generation module that generates answers related to the query and optimizes the generated content based on the prompt words;
[0044] A result optimization module for deduplicating the generated content, adjusting the result sorting, and ensuring that the answers are concise, accurate, and diverse.
[0045] Further, the prompt word optimization module is a Prompt module, including:
[0046] Chain of Thought (CoT) prompt: Gradually guide the model to complete the solution of complex problems and enhance traceability by making the reasoning process transparent;
[0047] Structured prompt: Clarify the task background and context to improve the model's understanding of the question semantics;
[0048] Example-driven prompt: Guide the model to generate high-quality results by providing examples of questions and answers;
[0049] Conditioned prompt: Control the style, format, and length of the generated answers according to the requirements.
[0050] Further, the generation module generates the final answer through the LLM, specifically including:
[0051] Semantic generation: Generate high-quality answers related to the question based on the retrieved context and optimized prompts.
[0052] Answer optimization: Polish the language and adjust the logic of the generated content to ensure that the output result is fluent and context-appropriate.
[0053] Answer customization: Adjust the presentation of the answer according to the conditioned prompt, including the writing style, format, and level of detail of the information.
[0054] The beneficial effects of the present invention are as follows: By combining efficient retrieval and generation, the accuracy and real-time performance of information processing are improved. The present invention includes collecting documents from multi-source data and performing preprocessing, generating high-dimensional vector representations through text chunking and vectorization, constructing a vocabulary index, and implementing hierarchical retrieval and keyword matching; at the same time, using question generation and question-question retrieval methods, combining with a large language model (LLM) to optimize the generated results, and further enhancing the reasoning ability of the large model through Chain of Thought (CoT) and the prompt word optimization module, ensuring that the generated answers are highly accurate and logical. The present invention can be widely applied to fields such as intelligent question answering, knowledge base construction, and real-time information retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] To more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0056] Figure 1 : A schematic flow chart of a knowledge base generation method based on RAG technology;
[0057] Figure 2 : A schematic flow chart of a document retrieval method based on RAG technology. Specific embodiments
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0059] As Figure 1 shown, a knowledge base generation method based on RAG technology is characterized by including the following steps:
[0060] S101, collect documents in different formats from diverse data sources in multiple professional fields.
[0061] The data sources should cover a wide range of knowledge backgrounds and domain information to ensure the comprehensiveness and diversity of the document content, providing a solid data foundation for subsequent analysis.
[0062] S102, document format conversion and preprocessing:
[0063] Format conversion, convert documents in different formats (PDF, Word, HTML, etc.) into text data uniformly;
[0064] Text preprocessing, perform semantic error correction, word segmentation, stop word removal, and noise data cleaning on the text data.
[0065] Ensure the clear data structure and accurate content, laying a foundation for subsequent vectorization and retrieval.
[0066] S103, cut the preprocessed text data, divide the large text into multiple logical segments, evaluate the response quality according to different cut sizes, find the best cut granularity, and use vectorization technology to convert the text content into high-dimensional vectors.
[0067] Ensure that the semantic information of each text chunk can be accurately captured to support semantic retrieval.
[0068] S104, Generate hypothetical questions. Generate hypothetical questions based on the chunked text to simulate the user's query scenario.
[0069] Through natural language processing technology (NLP) and preset templates, generate diverse question contexts to clarify the target questions for subsequent information retrieval. The core purpose of this step is to ensure that the retrieval results can accurately match the user's semantic needs.
[0070] S105, Build a vocabulary index. Build an efficient vocabulary index, including the following operations:
[0071] Inverted index. Associate keywords with the positions in the document content to achieve fast positioning.
[0072] Semantic index. Combine vector representations to support semantic similarity retrieval and improve the relevance of retrieval results.
[0073] S106, Layer the text library. Use text summarization technology to generate summaries for the documents; generate vectors and keywords for the full text content of the text.
[0074] S107, Evaluate the quality of retrieval results and dynamically adjust the processing flow.
[0075] Including:
[0076] Quality assessment. Use automated evaluation metrics and manual reviews to evaluate the relevance and accuracy of retrieval results.
[0077] Dynamic adjustment. According to the evaluation results, optimize the retrieval algorithm, vectorization method, or text chunking strategy to ensure that subsequent steps can provide higher-quality results.
[0078] Continuous optimization. Continuously feedback and adjust the text processing flow.
[0079] Such as Figure 2 As shown, a retrieval method based on the RAG technology includes the following steps:
[0080] S201, Vector retrieval. Convert the user's query semantics into a high-dimensional vector representation and compare it with the high-dimensional semantic vectors converted from the text data. Through similarity calculation, match the document set that is most relevant to the user's query semantics.
[0081] This step is applicable to solving queries with high semantic similarity but different surface vocabulary, ensuring the semantic matching accuracy of retrieval results.
[0082] S202, Keyword Retrieval, which matches the keywords in the user's query, quickly locates the documents containing these words, and filters the document set relevant to the user's query; for questions with weak semantic extensibility, further retrieval is combined with vectors.
[0083] Keyword retrieval includes:
[0084] Optimized keyword extraction, which uses advanced keyword extraction algorithms (such as TF-IDF, RAKE) to improve the accuracy of keyword matching;
[0085] Multi-keyword combination matching, which supports multi-keyword combination queries to cover a wider range of query requirements;
[0086] Keyword weight adjustment, which dynamically adjusts the keyword weights according to the importance of the document and the frequency of keyword occurrence to optimize the retrieval effect.
[0087] S203, Hierarchical Retrieval, which preliminarily screens the abstract part of the documents to quickly narrow down the range of candidate documents; among the documents screened by the abstract, in-depth analysis is carried out on the full text to mine the core information and detailed content;
[0088] S204, Question-Question Retrieval, which analyzes the context and keywords of the user's question. Combining the query context, the question is transformed into a structured retrieval statement; through the semantic matching model, the documents or answers relevant to the user's question are located; this step ensures the context consistency between the query and the retrieval results, making the system perform excellently in complex question answering scenarios.
[0089] S205, Rearrangement after Retrieval, which applies an advanced sorting model to perform a secondary sort on the preliminary results.
[0090] According to indicators such as the semantic relevance and confidence level between the document and the query, the sorting order is optimized to ensure that the most relevant results are displayed first, improving the user experience. The rearrangement process significantly improves the accuracy of result sorting and user satisfaction.
[0091] Rearrangement after retrieval includes:
[0092] Preliminary sorting, which uses traditional sorting algorithms to perform a preliminary sort on the candidate documents and generates a basic ranking list according to the matching degree between the query and the document;
[0093] Advanced sorting, which uses an advanced sorting model to optimize the preliminary sorting results and adjusts the priority of the documents by combining indicators such as semantic relevance, document authority, and content integrity;
[0094] Semantic relevance optimization, during the sorting process, the semantic relevance between the document and the query is re-evaluated by combining the context information of the user's query, and the most relevant and context-compliant document content is preferentially displayed;
[0095] Personalized sorting, which adjusts the sorting results according to the user's historical query records and preferences.
[0096] S206, Verify the retrieval results, and use an external verification system to check the consistency of the retrieval results.
[0097] Manually or automatically review the results of fuzzy queries or noise information, eliminate incorrect content, and ensure the accuracy and relevance of the output results. When the information complexity is high or the query is fuzzy, the verification step further improves the reliability of the results.
[0098] Verifying the retrieval results includes:
[0099] Use an external knowledge base or verification model to detect the semantic consistency of the retrieval results.
[0100] Evaluate the retrieved content through manual review and automated metrics (such as accuracy and recall), and eliminate irrelevant or noisy data.
[0101] S207, Deduplicate the information, automatically identify duplicate content through semantic similarity detection, remove redundant information, and reorganize and output according to the importance of the content.
[0102] Ensure that the output content is concise, clear, and meets the needs of diversification and practicality. The deduplication ability of the LLM further improves the readability and professionalism of the generated content.
[0103] An intelligent Q&A system based on RAG technology, including the following modules:
[0104] Document preprocessing module, used to collect documents from data sources in multiple professional fields and perform formatting and semantic cleaning;
[0105] Retrieval module, including vector retrieval, keyword retrieval, and hierarchical retrieval functions, used to quickly locate relevant documents;
[0106] Question analysis module, which analyzes the questions input by users, converts them into high-dimensional semantic vectors, and generates structured query statements;
[0107] Prompt optimization module, used to optimize the generation process.
[0108] The prompt optimization module is the Prompt module, including:
[0109] Chain reasoning prompt: Gradually guide the model to complete the answer to complex questions, and enhance traceability by transparentizing the reasoning process;
[0110] Structured prompt: Clearly define the task background and context, and improve the model's understanding of the question semantics;
[0111] Example-driven prompt: By providing examples of questions and answers, guide the model to generate high-quality results;
[0112] Conditional prompt: Control the style, format, and length of the generated answer according to requirements.
[0113] A generation module that generates answers related to the query and optimizes the generated content based on the prompt words;
[0114] The generation module generates the final answer through an LLM (Large Language Model), specifically including:
[0115] Semantic generation: Based on the retrieved context and optimized prompt, generate high-quality answers related to the question.
[0116] Answer optimization: Polish the language and adjust the logic of the generated content to ensure that the output result is fluent and contextually appropriate.
[0117] Answer customization: According to the conditional prompt, adjust the presentation of the answer, including the text style, format, and level of detail of the information.
[0118] A result optimization module used to deduplicate the generated content, adjust the result sorting, and ensure that the answers are concise, accurate, and diverse.
[0119] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation manners. Obviously, many modifications and variations can be made according to the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the present invention, so that those skilled in the relevant technical fields can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A knowledge base generation method based on RAG technology, characterized in that: The following steps are involved: S101, collects documents in different formats from diverse data sources in multiple professional fields; S102, document format conversion and preprocessing: Format conversion: convert documents in different formats (PDF, Word, HTML, etc.) into text data; Text preprocessing: semantic error correction, word segmentation, stop word removal, and noise data cleaning of text data; S103, slicing the preprocessed text data into multiple logical segments, evaluating the response quality according to different slicing sizes, finding the best slicing granularity, and using vectorization technology to convert the text content into a high-dimensional vector; S104, generating hypothetical questions, generating hypothetical questions based on the segmented text to simulate the user's query scenario; S105, vocabulary index construction, building an efficient vocabulary index, including the following operations: Inverted index, which associates keywords with the location of document content to achieve rapid positioning; Semantic indexing, combined with vector representation, supports semantic similarity retrieval and improves the relevance of retrieval results; S106, stratify the text library, generate a summary of the document using text summarization technology; generate vectors and keywords of the full text content; S107, evaluate the quality of the search results and dynamically adjust the processing flow.
2. The method for generating a knowledge base based on RAG technology according to claim 1, characterized in that: The step S107 includes: Quality assessment, using automated evaluation metrics and manual review to assess the relevance and accuracy of search results; Dynamic adjustment: Based on the evaluation results, the search algorithm, vectorization method or text segmentation strategy are optimized to ensure that subsequent steps can provide higher quality results. Continuous optimization, constant feedback and adjustment of text processing processes.
3. A retrieval method based on RAG technology, implemented based on the knowledge base generation method according to claim 1, characterized in that: The following specific methods are included: S201, vector retrieval, converts the user query semantics into a high-dimensional vector representation, and compares it with the high-dimensional semantic vector converted from the text data, and matches the document set most relevant to the user query semantics through similarity calculation; S202, keyword search, matching keywords in the user query, and quickly locating documents containing these words, filtering the document set related to the user query; for questions with weak semantic scalability, further search is performed in combination with vectors; S203, hierarchical retrieval, preliminary screening of the abstract part of the document, quickly narrowing down the candidate document range; among the documents screened by the abstract, in-depth analysis is performed on the full text to mine the core information and detailed content; S204, question-question retrieval, analyzes the context and keywords of the user's question. Combined with the query context, the question is converted into a structured search statement; through the semantic matching model, the document or answer related to the user's question is located; S205, re-ranking after retrieval, applying an advanced ranking model to perform secondary ranking on the preliminary results; S206, verifying the search results, using an external verification system to perform consistency check on the search results; S207, deduplicate the information, automatically identify duplicate content through semantic similarity detection, remove redundant information, and reorganize the output according to the importance of the content.
4. A retrieval method based on RAG technology according to claim 3, characterized in that: In step S202, the keyword search includes keyword extraction optimization, multi-keyword combination matching and keyword weight adjustment.
5. The retrieval method based on RAG technology according to claim 3, characterized in that: In S205, the post-search rearrangement includes: Preliminary sorting: Use traditional sorting algorithms to preliminarily sort candidate documents and generate a basic ranking list based on the matching degree between the query and the document; Advanced sorting: Use advanced sorting models to optimize the preliminary sorting results and adjust the document priority based on indicators such as semantic relevance, document authority, and content completeness; Semantic relevance optimization: During the sorting process, the semantic relevance between the document and the query is re-evaluated in combination with the contextual information of the user's query, and the most relevant and contextual document content is displayed first; Personalized sorting: personalize the sorting results based on the user's historical query records and preferences.
6. A retrieval method based on RAG technology according to claim 3, characterized in that: In step S206, verifying the search results includes: Use external knowledge base or verification model to check the semantic consistency of retrieval results. Retrieve content through manual review and automated indicator evaluation to eliminate irrelevant or noisy data.
7. An intelligent question-answering system based on RAG technology, used to implement the knowledge base generation method described in claim 1 and the retrieval method described in claim 3, characterized in that: Includes the following modules: Document preprocessing module, which is used to collect documents from data sources in multiple professional fields and perform formatting and semantic cleaning; Retrieval module, including vector retrieval, keyword retrieval and hierarchical retrieval functions, used to quickly locate relevant documents; The question analysis module analyzes the questions input by users, converts them into high-dimensional semantic vectors and generates structured query statements; Prompt word optimization module, used to optimize the generation process; The generation module generates answers relevant to the query and optimizes the generated content based on the prompt words; The result optimization module is used to deduplicate the generated content and adjust the result sorting to ensure that the answers are concise, accurate and diverse.
8. The intelligent question-answering system based on RAG technology according to claim 7, characterized in that: The prompt word optimization module is the Prompt module, including: Chained reasoning prompts: step-by-step guide the model to complete the solution of complex problems, enhancing traceability through transparent reasoning process; Structured prompts: clarify the task background and context to improve the model's understanding of the problem semantics; Example-driven prompts: Guide the model to produce high-quality results by providing examples of questions and answers; Conditional prompts: Control the style, format, and length of generated answers based on your needs.
9. The intelligent question-answering system based on RAG technology according to claim 7, characterized in that: The generation module generates the final answer through LLM, which specifically includes: Semantic generation generates high-quality answers relevant to the question based on the retrieval context and optimization hints. Optimize answers by polishing the language and adjusting the logic of generated content to ensure that the output is fluent and in context. Customize answers and adjust the way answers are presented based on conditional prompts, including text style, format, and details.
Citation Information
Cited By
Intention classification method and device based on vector retrieval and context awareness and medium
CN120448929A
Intent classification method, device and medium based on vector retrieval and context perception
CN120448929B
Contract retrieval enhancement optimization method, equipment and medium
CN120596647A
Low-altitude intelligent question and answer construction method and system based on dynamic parameters
CN120632055A
Financial consultation response method and system
CN120806166A