Knowledge base construction method and device supporting retrieval enhancement generation
By introducing concept layer and semantic layer into the RAG system, separating source data from the knowledge base, and performing concept filtering, error correction and diversified processing, the problems of low-quality and redundant content in the knowledge base are solved, and efficient and accurate knowledge base construction and retrieval are achieved.
Patent Information
- Application Number
- CN202510231559.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
When building a knowledge base, there is a large amount of low quality and redundant content, resulting in low retrieval efficiency and low accuracy.
By introducing concept layer and semantic layer, the source data and knowledge base are separated, the concept layer is used to extract concepts and filter, correct, enhance and diversify them through the semantic layer to generate knowledge bases that support retrieval enhancement generation.
Reduces data redundancy in the knowledge base, improves data accuracy and professionalism, enhances retrieval efficiency and accuracy, and supports efficient answers to complex questions.
Smart Images

Figure CN120146167A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language models and information retrieval, and particularly relates to a knowledge base construction method and device for supporting retrieval-augmented generation. Background Art
[0002] In recent years, large language models have demonstrated remarkable capabilities in various natural language processing tasks. However, when answering complex questions, large language models often fail to output answers with high accuracy due to the lack of the latest information. The RAG (Retrieval-augmented Generation) system can combine large language models and information retrieval. The RAG system retrieves relevant content in the knowledge base in real time to obtain the knowledge most relevant to the user's question, enabling the large language model to generate and output answers by combining the retrieved relevant background knowledge.
[0003] However, in existing RAG systems, the way to generate a knowledge base is usually as follows: the content of the source file is divided into several small segments, and then the small segments of text are converted into vectors through a text embedding model, and the converted vectors are stored in a vector database to obtain the knowledge base. Since there is a lot of content irrelevant to the substantial content of the file in many source files, storing all this low-quality content in the knowledge base will increase the storage and query costs of the database. The retrieval efficiency of the RAG system is low, and the accuracy of the retrieved answers is also low. Summary of the Invention
[0004] The purpose of the present invention is to separate the source data and the knowledge base, extract concepts from the source data through the concept layer, and process the concepts through the semantic layer, thereby reducing the data redundancy of the knowledge base, improving the data accuracy and professionalism of the knowledge base, and broadening the data content of the knowledge base, so as to improve the retrieval efficiency and retrieval accuracy of the retrieval-augmented generation system.
[0005] In a first aspect, an embodiment of the present invention provides a knowledge base construction method for supporting retrieval-augmented generation. The method is applied to a retrieval-augmented generation system, and the retrieval-augmented generation system includes a concept layer and a semantic layer. The method includes:
[0006] Obtain source data for constructing a knowledge base, and extract multiple concepts included in the source data through the concept layer;
[0007] Perform concept processing on the multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer; the concept processing includes one or more of the following processing methods: concept filtering, concept error correction, concept enhancement, and concept diversification;
[0008] Convert the concepts processed by the semantic layer into vectors, and store the obtained vectors in a vector database to obtain a knowledge base that supports retrieval-augmented generation;
[0009] Among them, the concept filtering is used to delete redundant concepts with repeated semantics, the concept correction is used to correct the concepts with semantic errors among the multiple concepts, the concept enhancement is used to generate composite concepts through the semantic information of the multiple concepts, and the concept diversification is used for: for any concept among the multiple concepts, based on the semantic information of the concept, generate multiple related concepts with different expression forms from the concept.
[0010] Optionally, the process of performing concept processing on the multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer includes:
[0011] Receive a deletion instruction for concepts with repeated semantics input by the user, and delete redundant concepts from the multiple concepts with repeated semantics based on the deletion instruction;
[0012] And / or,
[0013] Receive a correction instruction for concepts with semantic errors input by the user, and correct the concepts with semantic errors into concepts with correct semantics based on the correction instruction.
[0014] Optionally, the process of performing concept processing on the multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer includes:
[0015] Analyze the semantic information of each concept in the multiple concepts;
[0016] Based on the semantic information corresponding to the multiple concepts respectively, analyze the semantic repetition degree between different concepts in the multiple concepts;
[0017] When there is a case where the semantic repetition degree between multiple concepts is greater than a preset repetition degree, delete redundant concepts from the multiple concepts with repeated semantics.
[0018] Optionally, the process of performing concept processing on the multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer includes:
[0019] Analyze the semantic information of each concept in the multiple concepts;
[0020] Based on the semantic information of the multiple concepts, combine or associate multiple concepts with related semantic information to obtain composite concepts.
[0021] Optionally, the method further includes:
[0022] Determine the source data corresponding to each of the multiple concepts; wherein, the source data corresponding to each concept is used to characterize the initial source of the concept;
[0023] For each of the multiple concepts, associate the concept with the source data corresponding to the concept.
[0024] Optionally, after converting the concepts processed by the semantic layer into vectors and storing the obtained vectors in a vector database to obtain a knowledge base supporting retrieval-augmented generation, the method further includes:
[0025] Obtain a natural language question input by a user through a terminal;
[0026] Based on the natural language question, find the target vector corresponding to the target concept in the knowledge base whose relevance to the natural language question is greater than a preset relevance;
[0027] Generate a natural language answer based on the target vector corresponding to the target concept and the natural language question, and send the natural language answer to the terminal;
[0028] When receiving a preset instruction for the target concept in the natural language answer sent by the terminal, send the target source data corresponding to the target concept to the terminal, so that the content displayed on the display interface of the terminal jumps to the target source data; the preset instruction is used to indicate viewing the initial source data of the target concept.
[0029] In a second aspect, an embodiment of the present invention provides a knowledge base construction device for supporting retrieval-augmented generation. The device is applied to a retrieval-augmented generation system, and the retrieval-augmented generation system includes a concept layer and a semantic layer. The device includes:
[0030] A concept extraction module, configured to obtain source data for constructing a knowledge base and extract a plurality of concepts included in the source data through the concept layer;
[0031] A concept processing model, configured to perform concept processing on the plurality of concepts through the semantic layer to obtain concepts processed by the semantic layer; the concept processing includes one or more of the following processing methods: concept filtering, concept correction, concept enhancement, and concept diversification;
[0032] A knowledge base generation module, configured to convert the concepts processed by the semantic layer into vectors and store the obtained vectors in a vector database to obtain a knowledge base supporting retrieval-augmented generation;
[0033] Among them, the concept filtering is used to delete redundant concepts with repeated semantics, the concept error correction is used to correct the concepts with semantic errors among the multiple concepts, the concept enhancement is used to generate composite concepts through the semantic information of the multiple concepts, and the concept diversification is used to: for any concept among the multiple concepts, based on the semantic information of the concept, generate multiple related concepts with different expression forms from the concept.
[0034] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0035] At least one processor;
[0036] A memory for storing instructions executable by the at least one processor;
[0037] Among them, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.
[0038] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in the first aspect.
[0039] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect.
[0040] In the technical solution of the embodiment of the present invention, the retrieval enhancement generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept error correction, concept enhancement, and concept diversification, to obtain the concepts processed by the semantic layer; then the concepts processed by the semantic layer are converted into vectors, and the obtained vectors are stored in a vector database to obtain a knowledge base supporting retrieval enhancement generation.
[0041] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. Through the concept layer and the semantic layer, concepts are extracted from the source data and processed, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, generate concepts with more semantic information through concept enhancement, and through concept diversification, make the concepts have multiple expression forms. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, and broadens the data content of the knowledge base, which helps to improve the retrieval efficiency and retrieval accuracy of the retrieval enhancement generation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the overall technical solution provided by the embodiment of the present invention;
[0043] Figure 2 Flowchart of a knowledge base construction method for supporting retrieval-augmented generation provided by an embodiment of the present invention;
[0044] Figure 3 Structural schematic diagram of a knowledge base construction device for supporting retrieval-augmented generation provided by an embodiment of the present invention;
[0045] Figure 4 Structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0046] The present invention will be described in detail below through embodiments.
[0047] In recent years, large language models have demonstrated remarkable capabilities in various natural language processing tasks. However, when answering complex questions, large language models often fail to output answers with high accuracy due to the lack of the latest information. The Retrieval-augmented Generation (RAG) system can combine large language models with information retrieval. The RAG system retrieves relevant content in the knowledge base in real time to obtain the knowledge most relevant to the user's question, enabling the large language model to generate and output answers by combining the retrieved relevant background knowledge. In high-precision demand fields such as medical and legal fields, the application potential of the RAG system is particularly broad. It can support the dynamic answering of knowledge-intensive questions and is a key step in achieving interpretability and timeliness.
[0048] However, to build a knowledge base that supports retrieval-augmented generation, the source data needs to be parsed into a form that can be called by the large language model. Due to the limitations of the input content length of the text embedding model and the large language model, the current solutions for generating a knowledge base are usually as follows: First, the source data content needs to be divided into several small segments, and then the small segments of text are converted into vectors through the text embedding model and stored in the database to obtain the constructed knowledge base for easy retrieval and matching by the large language model. There are currently various methods for dividing the source data content into several small segments, which can be divided into the following four categories:
[0049] The first category: Fixed-length chunking method. Specifically, the source text is cut according to a fixed number of words or characters to ensure a unified chunk length. However, for semantically dense content, fixed-length chunking may cause content fragmentation.
[0050] The second category: Sliding window chunking method. Specifically, information is cross-overlapped in multiple chunks through an overlapping sliding window to enhance the correlation between chunks and ensure that key information is not easily lost during retrieval. At the same time, the data redundancy is also increased.
[0051] The third category: syntax-based chunking. Specifically, natural language processing tools are used to identify paragraphs, sentences, or chapters for chunking, which avoids information interruption, but the chunking efficiency for long paragraphs is low.
[0052] The fourth category: semantic-based chunking is a chunking method guided by content semantics. It uses natural language processing technology to automatically divide text into multiple small segments based on semantically similar content. Unlike simple fixed-length chunking, semantic chunking is based on the coherence and importance of contextual semantics. This ensures that each chunk has a relatively independent topic in the knowledge base, reduces the sense of fragmentation of the content, and improves the context coherence during retrieval. However, this method is highly complex.
[0053] As can be seen from the above description, in the related art, the method of dividing the source data content into several small segments is to divide the entire content of the source data content into several small segments. In this way, when generating the knowledge base, the several small segments are converted into vectors and stored in the knowledge base. Since there are a lot of content irrelevant to the actual content of the file in many source data, storing all these low-quality content in the knowledge base will increase the storage and query costs of the database, the retrieval efficiency of the RAG system is low, and the accuracy of the answers obtained by the retrieval is also low.
[0054] From the above description, it can be seen that the existing knowledge base construction technology mostly stores the source data directly into the knowledge base by partitioning the data into blocks. The block algorithm and the method of directly associating the source data with the knowledge base have the following shortcomings:
[0055] First, there is low-quality content in the knowledge base. Specifically, many source data contain a large amount of content that is irrelevant to the actual content of the document. Storing all of this low-quality content in the database will increase the storage and query costs of the database, and increase the noise of the database, resulting in context loss, which greatly increases the difficulty of subsequent retrieval algorithms, resulting in low retrieval efficiency and low accuracy of the answers retrieved.
[0056] Secondly, there is redundant content in the knowledge base. Specifically, in a specific professional field, for example, when the source data is medical literature and a database that supports retrieval enhancement generation is constructed through medical literature, there will be a lot of redundant content in the medical documents. Specifically, different medical literature will describe the research background and current status of the field in its overview and discussion chapters, resulting in a large amount of redundant content stored in the knowledge base. Similarly, these contents also increase the storage and query costs of the knowledge base, affecting the retrieval efficiency. At the same time, in top-k retrieval (that is, only the k most relevant answers are returned, k is a positive integer), a large number of similar answers may be ranked higher in the ranking, squeezing out other truly relevant and valuable answers, making it difficult for the retrieval results to focus on the core answers required by users, which leads to low retrieval accuracy.
[0057] Thirdly, in addition, the way of directly associating source data with the knowledge base affects the independence of the knowledge base and the flexibility of knowledge base update and maintenance. Especially for technical personnel, updating the knowledge base requires reprocessing and embedding of the source data in the old and new versions, consuming a large amount of resources and time. This may also lead to a decrease in the response speed of the knowledge base and a low retrieval efficiency, affecting the real-time experience of users.
[0058] In summary, the existing database construction method for supporting retrieval-augmented generation will increase the storage and query costs of the database. The RAG system has a low retrieval efficiency, and the accuracy of the answers obtained by retrieval is also low.
[0059] In view of the above technical problems existing in the related technologies, the retrieval-augmented generation system of the present invention proposes a solution of a separated architecture for source data and the knowledge base. The core of the separated architecture lies in introducing a concept layer and a semantic layer to achieve more efficient and flexible knowledge base management, so as to fundamentally solve the above technical problems.
[0060] For the clarity of the solution, first, in combination with Figure 1 , the overall technical solution of the embodiments of the present invention will be elaborated in detail.
[0061] (1) For the concept layer, it mainly includes the following aspects:
[0062] I. Obtain source data. For the medical profession, text data including clinical guidelines, medical literature, etc. can be obtained as source data, and through a concept extraction algorithm, it is converted into a series of traceable concepts.
[0063] By converting the text content of the source data into a conceptual expression, the concept layer not only extracts the key information of the source data but also avoids redundant content and irrelevant data from entering the knowledge base. For example, in medical literature, the common research background and literature review sections contain a large amount of repetitive descriptions. Through concept extraction, the focus can be on the key findings, conclusions, and unique information of the literature, thus avoiding redundant content from appearing in the knowledge base. In addition, concept extraction can also exclude misleading content (such as background information) and irrelevant noise data, thereby improving the retrieval accuracy and retrieval efficiency of the retrieval-augmented generation system.
[0064] II. The concept layer can provide a traceability function for each concept. Specifically, the concept layer can associate each concept with its original source data, so as to ensure the transparency and reliability of the knowledge base. Through the traceability of concepts, users can jump from a concept to the corresponding source data to further view the context and origin of the concept. This not only enhances the credibility of the knowledge base but also facilitates its use in high-standard applications such as scientific research and medicine, and is more conducive to the review and verification of source data.
[0065] III. Knowledge base management for professionals. Specifically, the concept extraction process does not require the support of IT personnel. Users in professional fields can manually maintain and update the concepts in the concept layer according to their needs. In this way, experts in fields such as medicine and scientific research can directly update the concept library in the concept layer based on the latest discoveries or changes, without the need for technicians to re-embed or adjust the knowledge base. This flexibility accelerates the update speed of the knowledge base, ensures the forefront and practicality of the knowledge base content, and thus can improve the retrieval accuracy and efficiency of the retrieval enhanced generation system.
[0066] (2) For the semantic layer, it mainly includes the following aspects:
[0067] I. Providing a verification function for professionals: The introduction of the semantic layer provides a function for manual screening and verification of concepts. Through the semantic layer, the automatically extracted concepts can be manually filtered and reviewed to ensure the accuracy of the concepts and compliance with the requirements of professional knowledge. Manual intervention can not only improve the quality of data but also correct errors or inconsistencies in the automated process, ensuring the credibility of the extracted concepts.
[0068] Specifically, among the multiple concepts automatically extracted by the concept layer, there may be concepts with semantic repetition, the probability of semantic errors, and concepts with non-standard expressions for professional fields.
[0069] The retrieval enhanced generation system can receive a deletion instruction from a professional for a concept with semantic repetition and can delete the redundant concept among the concepts with semantic repetition according to the instruction of the deletion instruction. For example, if among the multiple concepts extracted by the concept layer, there are three concepts with semantic repetition, namely concept 1, concept 2, and concept 3, then concept 2 and 3 can be deleted, and only concept 1 is retained, which can reduce concept redundancy and help reduce data redundancy in the knowledge base.
[0070] The retrieval enhanced generation system can also receive a correction instruction from a professional for a concept with a semantic error and correct the concept with a semantic error to a concept with a correct semantics according to the correction instruction, which can then improve the accuracy of the concept and help improve the retrieval accuracy.
[0071] Similarly, the retrieval enhanced generation system can also receive a diversification expression instruction from a professional for any concept among multiple concepts and generate multiple different expression forms according to these instructions. This diversification processing can not only enrich the expression forms of concepts but also improve the recall rate of information retrieval, enabling the system to more comprehensively capture the user's query intent. By realizing the diversification of concepts at the semantic level, the data in the knowledge base will be more flexible and diverse, facilitating the satisfaction of different users' needs and understanding.
[0072] Specifically, for any concept, multiple expressions of the concept can be enriched and expanded through synonyms, related questions, or variants, thereby obtaining multiple related concepts with different expressions from the concept. For example, for the concept "Apples are red", multiple related concepts with different expressions such as "What is the color of apples?" or "What are the red fruits?" can be generated. Thus, when different users query, even if the natural language questions entered have different expressions, the response ability of the retrieval augmented generation system is relatively high, thereby improving the user experience.
[0073] II. Concept enhancement. By analyzing the semantic information of different concepts, related concepts with relevance can be combined or associated, thereby generating more hierarchical and in-depth composite concepts. This concept enhancement helps to analyze and understand natural language questions from multiple dimensions, provides a richer information background, makes the provided natural language answers more accurate and intelligent, and can support more complex natural language question analysis and decision-making.
[0074] III. Automatically removing redundant concepts at the concept layer. An important function of the semantic layer is to remove redundant concepts with semantic repetition. By analyzing the semantic information of different concepts and the semantic information repetition degree of different concepts, if the semantic repetition degree between different concepts is relatively high, the redundant concepts with semantic repetition can be automatically removed, thereby reducing concept redundancy, helping to reduce data redundancy in the knowledge base, optimizing data storage, and improving the efficiency of data query and retrieval.
[0075] IV. The semantic layer provides an abstraction layer for concepts and can make the large language model more flexible and capable of answering complex questions through semantic extension and update. With the addition of new concepts, the semantic layer can be continuously updated and improved to ensure the long-term stability and scalability of the retrieval augmented generation system.
[0076] Through the above four aspects, the semantic layer processes the concepts extracted by the concept layer to obtain the concepts after being processed by the semantic layer, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, and generate concepts with more semantic information through concept enhancement.
[0077] (3) Next, the concepts processed by the semantic layer are transformed into vectors through a text embedding model and stored in a vector database to obtain a knowledge base that supports retrieval augmented generation.
[0078] Among them, the process of converting the concepts processed at the semantic layer into vectors through a text embedding model and the process of storing the vectors in a vector database are both fully automated. All newly generated or updated concepts can be automatically converted into vectors and stored in the vector database. This automated process decouples the update of the knowledge base data from the update of the vector database, simplifies the technical process, and reduces the complexity of system update and maintenance. The knowledge base constructed with this framework serves as the data support for the retrieval augmented generation system, enabling the large semantic model to provide more accurate retrieval results. For example, in the medical field, it helps to provide more accurate disease descriptions, medical analyses, and treatment suggestions.
[0079] After constructing a knowledge base that supports retrieval augmented generation, the retrieval augmented generation system can achieve high-efficiency and high-accuracy retrieval. The specific process can be as follows: Obtain the natural language question input by the user through the terminal, convert the natural language question into a vector, and search in the knowledge base for the target vector corresponding to the target concept whose vector relevance to the natural language question is greater than the preset relevance. Finally, generate a natural language answer based on the target vector corresponding to the target concept and the natural language question, and send the natural language answer to the terminal.
[0080] In the technical solution of the embodiment of the present invention, the retrieval augmented generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept correction, concept enhancement, and concept diversification to obtain the concepts processed by the semantic layer; then convert the concepts processed by the semantic layer into vectors, and store the converted vectors in a vector database to obtain a knowledge base that supports retrieval augmented generation.
[0081] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. By extracting concepts from the source data through the concept layer and the semantic layer and processing the concepts, it helps to reduce concept redundancy, improve the accuracy and professionalism of the concepts, generate concepts with more semantic information through concept enhancement, and make the concepts have multiple expressions through concept diversification. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, and broadens the data content of the knowledge base, which helps to improve the retrieval efficiency and retrieval accuracy of the retrieval augmented generation system.
[0082] In addition, the concept layer also provides a manual verification function, which further improves the accuracy and professionalism of the concepts. Moreover, the concept layer and the semantic layer support reverse knowledge traceability. When the large language model answers complex questions, it can provide the source data that supports its point of view, thereby improving the accuracy of the answers to the questions. At the same time, the generated knowledge base uses a vector library rather than a graph database structure, ensuring that the retrieval enhancement generation system is highly flexible, scalable, and low in complexity. By separating the source data and the vector database, it can ensure that the update and maintenance of the retrieval enhancement generation system does not require complex IT knowledge, which is conducive to professionals to update the knowledge base.
[0083] After the overall technical solution of the embodiment of the present invention is described in detail, a knowledge base construction method supporting retrieval enhanced generation provided by the embodiment of the present invention will be described in detail below.
[0084] The embodiment of the present invention provides a knowledge base construction method supporting retrieval enhancement generation, the method is applied to a retrieval enhancement generation system, the retrieval enhancement generation system includes a concept layer and a semantic layer, such as Figure 2 As shown, the method may include the following steps:
[0085] S210, acquiring source data for building a knowledge base, and extracting multiple concepts included in the source data through a concept layer.
[0086] Specifically, for the medical profession, text data including clinical guidelines and medical literature can be obtained as source data, and converted into a series of traceable concepts through concept extraction algorithms.
[0087] The concept layer converts the text content of the source data into conceptual expressions, which not only extracts the key information of the source data, but also prevents redundant content and irrelevant data from entering the knowledge base. For example, in medical literature, the common research background and literature review sections contain a large number of repetitive descriptions. Through concept extraction, we can focus on the key findings, conclusions, and unique information of the literature, thereby avoiding redundant content from appearing in the knowledge base. In addition, concept extraction can also exclude misleading content (such as background information) and irrelevant noise data, thereby improving the retrieval accuracy and efficiency of the retrieval enhancement generation system.
[0088] In addition, the concept extraction process does not require the support of IT personnel, and users in professional fields can manually maintain and update the concepts in the concept layer according to their needs. In this way, experts in fields such as medicine and scientific research can directly update the concept library in the concept layer based on the latest discoveries or changes, without the need for technical personnel to re-embed or adjust the knowledge base. This flexibility accelerates the update speed of the knowledge base, ensures the cutting-edge and practical nature of the knowledge base content, and thus improves the retrieval accuracy and efficiency of the retrieval enhancement generation system.
[0089] S220. Concept processing is performed on multiple concepts through the semantic layer to obtain the concepts after semantic layer processing. The concept processing includes one or more of the following processing methods: concept filtering, concept error correction, concept enhancement, and concept diversification.
[0090] Among them, concept filtering is used to delete redundant concepts with repeated semantics, concept error correction is used to correct the concepts with semantic errors among multiple concepts, concept enhancement is used to generate composite concepts through the semantic information of multiple concepts, and concept diversification is used for: for any one of the multiple concepts, based on the semantic information of the concept, generate multiple related concepts with different expression forms from the concept.
[0091] Specifically, since the concept layer is multiple concepts extracted by a concept extraction algorithm, there may be redundant concepts, incorrect concepts, or concepts that are not standardized for a professional field among the multiple concepts extracted by the concept layer; and the multiple concepts are independent of each other and not combined or associated together. Therefore, concept processing is performed on multiple concepts through the semantic layer to obtain the concepts after semantic layer processing. Among them, the concept processing may include one or more of the following processing methods: concept filtering, concept error correction, concept enhancement, and concept diversification.
[0092] The probability of redundant concepts, incorrect concepts, or non-professional concepts existing in the concepts after semantic layer processing is relatively small, and the concept expressions are diversified. That is, processing the concepts through the semantic layer helps to reduce concept redundancy, improve the accuracy and professionalism of the concepts, generate concepts with more semantic information through concept enhancement, and make the concepts have multiple expression forms through concept diversification.
[0093] For the sake of clear description of the solution, the specific implementation of "S220. Concept processing is performed on multiple concepts through the semantic layer to obtain the concepts after semantic layer processing" will be elaborated in detail in the following embodiments.
[0094] S230. Convert the concepts after semantic layer processing into vectors, and store the converted vectors in a vector database to obtain a knowledge base that supports retrieval-enhanced generation.
[0095] Specifically, after obtaining the concepts after semantic layer processing, the concepts after semantic layer processing can be input into a text embedding model, and the concepts after semantic layer processing are converted into vectors through the text embedding model, and the converted vectors are stored in a vector database to obtain a knowledge base that supports retrieval-enhanced generation.
[0096] The process of converting the concepts processed at the semantic layer into vectors through a text embedding model and the process of storing the vectors in a vector database are both completely automated. All newly generated or updated concepts can be automatically converted into vectors and stored in the vector database. This automated process decouples the update of the knowledge base data from the update of the vector database, simplifies the technical process, and reduces the complexity of system update and maintenance. The knowledge base constructed with this framework serves as the data support for the retrieval augmented generation system, enabling the large semantic model to provide more accurate retrieval results. For example, in the medical field, it helps to provide more accurate disease descriptions, medical analyses, and treatment suggestions.
[0097] In the technical solution of the embodiment of the present invention, the retrieval augmented generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept error correction, concept enhancement, and concept diversification to obtain the concepts processed by the semantic layer. Then, the concepts processed by the semantic layer are converted into vectors, and the obtained vectors are stored in a vector database to obtain a knowledge base that supports retrieval augmented generation.
[0098] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. Through the concept layer and the semantic layer, concepts are extracted from the source data and processed, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, generate concepts with more semantic information through concept enhancement, and make concepts have multiple expressions through concept diversification. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, broadens the data content of the knowledge base, and helps to improve the retrieval efficiency and retrieval accuracy of the retrieval augmented generation system.
[0099] For the sake of clear description of the solution, the specific implementation of "S220, performing concept processing on multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer" will be elaborated in detail below.
[0100] In one implementation, S220, performing concept processing on multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer, may include the following step a:
[0101] Step a, receiving a deletion instruction for concepts with semantic repetition input by the user, and deleting redundant concepts from the multiple concepts with semantic repetition based on the deletion instruction.
[0102] Specifically, the introduction of the semantic layer provides a function for manual screening of concepts, that is, professionals can perform manual screening on concepts. Through the semantic layer, the automatically extracted concepts can be manually filtered to delete redundant concepts.
[0103] The retrieval enhanced generation system can receive deletion instructions from professionals for semantically duplicate concepts and delete redundant concepts among the semantically duplicate concepts according to the instructions of the deletion instructions. For example, if among multiple concepts extracted at the concept layer, there are three semantically duplicate concepts, namely Concept 1, Concept 2, and Concept 3, then Concept 2 and 3 can be deleted, and only Concept 1 is retained, thereby reducing concept redundancy and helping to reduce data redundancy in the knowledge base.
[0104] In another implementation, S220, concept processing is performed on multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer, which may include the following step b:
[0105] Step b: Receive a correction instruction from the user for a concept with semantic errors and correct the concept with semantic errors into a concept with correct semantics based on the correction instruction.
[0106] Specifically, the introduction of the semantic layer provides a function for manually verifying concepts. Through the semantic layer, the automatically extracted concepts can be manually reviewed to ensure the accuracy of the concepts. Manual intervention can not only improve the quality of data but also correct errors in the automation process and ensure the credibility of the extracted concepts.
[0107] The retrieval enhanced generation system can also receive a correction instruction from a professional for a concept with semantic errors and correct the concept with semantic errors into a concept with correct semantics according to the correction instruction, thereby improving the accuracy of the concept and helping to improve the retrieval accuracy.
[0108] In another implementation, S220, concept processing is performed on multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer, which may include the following steps, namely step c1 to step c3:
[0109] Step c1: Analyze the semantic information of each concept among multiple concepts.
[0110] Step c2: Based on the semantic information corresponding to multiple concepts respectively, analyze the semantic duplication degree between different concepts among multiple concepts.
[0111] Step c3: When there is a case where the semantic duplication degree between multiple concepts is greater than the preset duplication degree, delete redundant concepts from the multiple semantically duplicate concepts.
[0112] An important function of the semantic layer is to remove redundant concepts with semantic duplication. By analyzing the semantic information of different concepts and analyzing the semantic information duplication degree of different concepts, if the semantic duplication degree between different concepts is relatively high, the redundant concepts with semantic duplication can be automatically removed, thereby reducing concept redundancy, helping to reduce data redundancy in the knowledge base, optimizing data storage, and improving the efficiency of data query and retrieval.
[0113] In another implementation, in S220, concept processing is performed on multiple concepts through a semantic layer to obtain the concepts after semantic layer processing, which may include the following steps, namely step d1 and step d2 respectively:
[0114] Step d1, analyze the semantic information of each concept in the multiple concepts;
[0115] Step d2, based on the semantic information of the multiple concepts, combine or associate the multiple concepts with related semantic information to obtain a composite concept.
[0116] Specifically, by analyzing the semantic information of different concepts, different concepts with a certain degree of association can be combined or associated, thereby generating a more hierarchical and in-depth composite concept. This concept enhancement helps to analyze and understand natural language problems from multiple dimensions, provides a richer information background, makes the provided natural language answers more accurate and intelligent, and can support more complex natural language problem analysis and decision-making.
[0117] Based on the above embodiments, in one implementation, the method for constructing a knowledge base that supports retrieval-enhanced generation may further include the following steps, namely step e1 and step e2 respectively:
[0118] Step e1, determine the source data corresponding to each concept in the multiple concepts; wherein, the source data corresponding to each concept is used to represent the initial source of the concept.
[0119] Step e2, for each concept in the multiple concepts, associate the concept with the source data corresponding to the concept.
[0120] Specifically, the concept layer can provide a traceability function for each concept. The concept layer can associate each concept with its original source data, so as to ensure the transparency and reliability of the knowledge base. Through the traceability of concepts, users can jump from a concept to the corresponding source data to further view the context and origin of the concept. This not only enhances the credibility of the knowledge base, but also facilitates its use in high-standard applications such as scientific research and medicine, and is more conducive to the review and verification of source data.
[0121] Based on the above embodiments, in one implementation, after converting the concepts after semantic layer processing into vectors and storing the obtained vectors in a vector database to obtain a knowledge base that supports retrieval-enhanced generation, the method for constructing a knowledge base that supports retrieval-enhanced generation may further include the following steps, namely step f1 to step f4:
[0122] Step f1, obtain the natural language question input by the user through the terminal.
[0123] Step f2: Based on the natural language question, search the knowledge base for the target vector corresponding to the target concept whose relevance to the natural language question is greater than a preset relevance threshold.
[0124] Step f3: Generate a natural language answer based on the target vector corresponding to the target concept and the natural language question, and send the natural language answer to the terminal.
[0125] Step f4: When receiving a preset instruction for the target concept in the natural language answer sent by the terminal, send the target source data corresponding to the target concept to the terminal, so that the content displayed on the terminal display interface jumps to the target source data.
[0126] Wherein, the preset instruction is used to indicate viewing the initial source data of the target concept.
[0127] Specifically, after constructing a knowledge base that supports retrieval-augmented generation, the retrieval-augmented generation system can achieve high-efficiency and high-accuracy retrieval. The specific process can be as follows: Obtain the natural language question input by the user through the terminal, convert the natural language question into a vector, search the knowledge base for the target vector corresponding to the target concept whose relevance to the vector corresponding to the natural language question is greater than a preset relevance threshold, and finally generate a natural language answer based on the target vector corresponding to the target concept and the natural language question, and send the natural language answer to the terminal. Since the data redundancy in the knowledge base is low, the data accuracy and professionalism in the knowledge base are high, and the data content in the knowledge base is relatively rich, the retrieval efficiency and retrieval accuracy of the retrieval-augmented generation system are both high.
[0128] Moreover, when receiving an instruction to view the source data source of the target concept in the natural language answer sent by the terminal, the retrieval-augmented generation system can send the target source data corresponding to the target concept to the terminal. After receiving the target source data, the terminal can jump the content displayed on the display interface to the target source data. By tracing the origin of the concept, the user can jump from the concept to the corresponding source data to further view the context and origin of the concept. This not only enhances the credibility of the knowledge base, but also facilitates its use in high-standard applications such as scientific research and medicine, and is more conducive to the review and verification of the source data.
[0129] An embodiment of the present invention further provides a device for constructing a knowledge base that supports retrieval-augmented generation. The device is applied to a retrieval-augmented generation system, and the retrieval-augmented generation system includes a concept layer and a semantic layer, as Figure 3 shown. The device includes:
[0130] A concept extraction module 310, configured to obtain source data for constructing the knowledge base, and extract multiple concepts included in the source data through the concept layer;
[0131] A concept processing model 320 is used to perform concept processing on the multiple concepts through the semantic layer to obtain the concepts processed by the semantic layer. The concept processing includes one or more of the following processing methods: concept filtering, concept error correction, concept enhancement, and concept diversification.
[0132] A knowledge base generation module 330 is used to convert the concepts processed by the semantic layer into vectors and store the converted vectors in a vector database to obtain a knowledge base that supports retrieval enhancement generation.
[0133] Among them, the concept filtering is used to delete redundant concepts with repeated semantics, the concept error correction is used to correct the concepts with semantic errors in the multiple concepts, the concept enhancement is used to generate composite concepts through the semantic information of the multiple concepts, and the concept diversification is used for: for any concept in the multiple concepts, based on the semantic information of the concept, generate multiple related concepts with different expression forms from the concept.
[0134] In the technical solution of the embodiment of the present invention, the retrieval enhancement generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept error correction, concept enhancement, and concept diversification to obtain the concepts processed by the semantic layer. Then, the concepts processed by the semantic layer are converted into vectors, and the converted vectors are stored in a vector database to obtain a knowledge base that supports retrieval enhancement generation.
[0135] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. Through the concept layer and the semantic layer, concepts are extracted from the source data and processed, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, generate concepts with more semantic information through concept enhancement, and through concept diversification, make the concepts have multiple expression forms. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, and broadens the data content of the knowledge base, which helps to improve the retrieval efficiency and retrieval accuracy of the retrieval enhancement generation system.
[0136] In a third aspect, an embodiment of the present invention provides an electronic device, as Figure 4 shown, including:
[0137] At least one processor 401;
[0138] A memory 402 for storing instructions executable by the at least one processor;
[0139] Among them, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.
[0140] In the technical solution of the embodiment of the present invention, the retrieval enhanced generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept correction, concept enhancement, and concept diversification to obtain the concepts processed by the semantic layer; then the concepts processed by the semantic layer are converted into vectors, and the obtained vectors are stored in the vector database to obtain a knowledge base that supports retrieval enhanced generation.
[0141] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. Through the concept layer and the semantic layer, concepts are extracted from the source data and processed, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, and generate concepts with more semantic information through concept enhancement. Moreover, through concept diversification, concepts have multiple expression forms. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, and broadens the data content of the knowledge base, which helps to improve the retrieval efficiency and retrieval accuracy of the retrieval enhanced generation system.
[0142] In a fourth aspect, the embodiment of the present invention provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device can execute the method described in the first aspect.
[0143] In the technical solution of the embodiment of the present invention, the retrieval enhanced generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept correction, concept enhancement, and concept diversification to obtain the concepts processed by the semantic layer; then the concepts processed by the semantic layer are converted into vectors, and the obtained vectors are stored in the vector database to obtain a knowledge base that supports retrieval enhanced generation.
[0144] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. Through the concept layer and the semantic layer, concepts are extracted from the source data and processed, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, and generate concepts with more semantic information through concept enhancement. Moreover, through concept diversification, concepts have multiple expression forms. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, and broadens the data content of the knowledge base, which helps to improve the retrieval efficiency and retrieval accuracy of the retrieval enhanced generation system.
[0145] In a fifth aspect, the embodiment of the present invention provides a computer program product, including a computer program, and the computer program implements the method described in the first aspect when executed by a processor.
[0146] In the technical solution of the embodiment of the present invention, the retrieval enhanced generation system adds two data layers, namely the concept layer and the semantic layer. The concept layer is used to extract concepts from the source data, and the semantic layer is used to perform concept processing on the concepts extracted by the concept layer. The concept processing mainly includes concept filtering, concept error correction, concept enhancement, and concept diversification to obtain the concepts processed by the semantic layer. Then, the concepts processed by the semantic layer are converted into vectors, and the obtained vectors are stored in the vector database to obtain a knowledge base that supports retrieval enhanced generation.
[0147] It can be seen that in the embodiment of the present invention, the source data and the knowledge base are separated. Through the concept layer and the semantic layer, concepts are extracted from the source data and processed, which helps to reduce concept redundancy, improve the accuracy and professionalism of concepts, generate concepts with more semantic information through concept enhancement, and make concepts have multiple expression forms through concept diversification. Furthermore, it reduces the data redundancy of the knowledge base, improves the data accuracy and professionalism of the knowledge base, broadens the data content of the knowledge base, and helps to improve the retrieval efficiency and retrieval accuracy of the retrieval enhanced generation system.
[0148] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and purposes of the present invention.
Claims
1. A method for constructing a knowledge base supporting retrieval-enhanced generation, characterized in that: The method is applied to a search enhancement generation system, the search enhancement generation system includes a concept layer and a semantic layer, and the method includes: Acquire source data for building a knowledge base, and extract multiple concepts included in the source data through the concept layer; Performing concept processing on the multiple concepts through the semantic layer to obtain concepts processed by the semantic layer; the concept processing includes one or more of the following processing methods: concept filtering, concept error correction, concept enhancement and concept diversification; Converting the concepts processed by the semantic layer into vectors, and storing the converted vectors into a vector database to obtain a knowledge base supporting retrieval enhancement generation; Among them, the concept filtering is used to delete redundant concepts with repeated semantics, the concept correction is used to correct semantically erroneous concepts among the multiple concepts, the concept enhancement is used to generate compound concepts through the semantic information of the multiple concepts, and the concept diversification is used to: for any concept among the multiple concepts, based on the semantic information of the concept, generate multiple related concepts that are different from the concept expression method.
2. The method according to claim 1, characterized in that: The performing conceptual processing on the multiple concepts through the semantic layer to obtain concepts processed by the semantic layer includes: receiving a deletion instruction for a semantically repeated concept input by a user, and deleting a redundant concept from a plurality of semantically repeated concepts based on the deletion instruction; and / or, A correction instruction for a semantically incorrect concept input by a user is received, and the semantically incorrect concept is corrected into a semantically correct concept based on the correction instruction.
3. The method according to claim 1, characterized in that The performing conceptual processing on the multiple concepts through the semantic layer to obtain concepts processed by the semantic layer includes: analyzing semantic information of each of the plurality of concepts; Analyzing semantic repetition between different concepts in the multiple concepts based on semantic information corresponding to the multiple concepts respectively; When the semantic repetition degree between multiple concepts is greater than a preset repetition degree, redundant concepts are deleted from the multiple semantically repeated concepts.
4. The method according to claim 1, characterized in that: The performing conceptual processing on the multiple concepts through the semantic layer to obtain concepts processed by the semantic layer includes: analyzing semantic information of each of the plurality of concepts; Based on the semantic information of the multiple concepts, multiple concepts with related semantic information are combined or associated to obtain a composite concept.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Determine source data corresponding to each of the plurality of concepts; wherein the source data corresponding to each concept is used to characterize an initial source of the concept; For each concept of the plurality of concepts, the concept is associated with source data corresponding to the concept.
6. The method according to claim 5, characterized in that After converting the concepts processed by the semantic layer into vectors and storing the converted vectors in a vector database to obtain a knowledge base supporting retrieval enhancement generation, the method further includes: Obtain natural language questions input by users through the terminal; Based on the natural language question, searching from the knowledge base for a target vector corresponding to a target concept whose relevance to the natural language question is greater than a preset relevance; Generate a natural language answer based on the target vector corresponding to the target concept and the natural language question, and send the natural language answer to the terminal; When receiving a preset instruction sent by the terminal for the target concept in the natural language answer, the target source data corresponding to the target concept is sent to the terminal so that the content displayed in the terminal display interface jumps to the target source data; the preset instruction is used to instruct to view the initial source data of the target concept.
7. A knowledge base construction device supporting retrieval enhanced generation, characterized in that: The device is applied to a search enhancement generation system, the search enhancement generation system includes a concept layer and a semantic layer, and the device includes: A concept extraction module, used to obtain source data for building a knowledge base, and extract multiple concepts included in the source data through the concept layer; A concept processing model, used to perform concept processing on the multiple concepts through the semantic layer to obtain concepts processed by the semantic layer; the concept processing includes one or more of the following processing methods: concept filtering, concept error correction, concept enhancement and concept diversification; A knowledge base generation module, used to convert the concepts processed by the semantic layer into vectors, and store the converted vectors into a vector database to obtain a knowledge base supporting retrieval enhancement generation; Among them, the concept filtering is used to delete redundant concepts with repeated semantics, the concept correction is used to correct semantically erroneous concepts among the multiple concepts, the concept enhancement is used to generate compound concepts through the semantic information of the multiple concepts, and the concept diversification is used to: for any concept among the multiple concepts, based on the semantic information of the concept, generate multiple related concepts that are different from the concept expression method.
8. An electronic device, characterized in that: include: at least one processor; a memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.
Citation Information
Cited By
Retrieval enhancement generation method suitable for building structure field
CN121301589A