Intelligent question answering enhancement method and system based on RAG technology
By combining multi-path recall and keyword knowledge graphs, the problem of incomplete or redundant information in intelligent question answering systems is solved, and comprehensive and reliable answer results are achieved.
Patent Information
- Application Number
- CN202511213649.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing intelligent question-answering systems based on RAG technology struggle to fully cover the required information when using a single recall method, resulting in incomplete answers or issues of information redundancy and omissions.
A multi-path recall method is adopted, including keyword recall, semantic similarity recall, and text summary recall. A keyword knowledge graph is constructed, and the keyword usage rate in the answer results is used for optimization and supplementation. Finally, the final answer results are constructed.
By combining multi-path recall and keyword knowledge graphs, the intelligent question answering system can comprehensively cover potentially relevant text blocks, deeply explore logical relationships and semantic connections, and provide complete and reliable answer results.
Smart Images

Figure CN120723877B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to an intelligent question-answering enhancement method and system based on RAG technology. Background Technology
[0002] With the development of artificial intelligence technology, intelligent question answering systems based on retrieval augmented generation (RAG) have become a core solution for information retrieval and generation tasks. Traditional RAG systems effectively alleviate the problems of knowledge solidification and illusion in large models through the paradigm of vectorized retrieval recall + large model generation.
[0003] However, the reliability of intelligent question answering results based on RAG technology is mainly determined by the vectorized retrieval results. There are multiple existing retrieval methods, such as keyword retrieval, semantic similarity retrieval, and text summary retrieval. Each retrieval method has different advantages and disadvantages. Using only one retrieval method is difficult to retrieve the required text information, resulting in incomplete answers. If multiple retrieval methods are used simultaneously, the answers may be concise, redundant, or incomplete. Summary of the Invention
[0004] To address the aforementioned problems, this invention provides an intelligent question-answering enhancement method and system based on RAG technology.
[0005] The intelligent question-answering enhancement method and system based on RAG technology of the present invention adopts the following technical solution:
[0006] One embodiment of the present invention provides an intelligent question-answering enhancement method based on RAG technology, the method comprising the following steps:
[0007] A vector database for constructing a knowledge base is used to perform multi-path recall based on user questions. The multi-path recall includes at least: keyword recall, semantic similarity recall, and text summary recall. Each recall result contains several text blocks retrieved from the vector database, and each text block in each recall result corresponds to a recall index.
[0008] The combined recall results from multiple recall paths are input into a large model to obtain the first answer result. The keyword knowledge graph is constructed using the usage rate of keywords in the combined recall results from multiple recall paths in the first answer result. The second answer result is obtained by inputting the recall results from each path into the large model. The keyword knowledge graph is supplemented using the usage rate of keywords in the second answer result from each path.
[0009] The recall metrics for each recall result are updated based on the supplemented keyword knowledge graph. Using the updated recall metrics, several text blocks are re-recalled from the common recall results of multiple recalls. These text blocks are then input into the large model to obtain the final answer result.
[0010] The keyword knowledge graph is supplemented by using the usage rate of keywords in several text blocks in the final answer results. The user questions and the supplemented keyword knowledge graph are then stored in a vector database.
[0011] Preferably, the specific steps for obtaining the usage rate of keywords in the first response result from the combined recall results of the multi-path recall are as follows:
[0012] For any keyword in the common recall results, and for all statements in the first response results; encode each statement in the first response results into a statement vector, encode the user question into a question vector, calculate the cosine similarity between the statement vector and the question vector, and record it as the semantic feature of each statement; record all statements containing the keyword as target statements, and record the cosine similarity between the statement vector and the question vector corresponding to each target statement as the target feature of each target statement.
[0013] Clustering the semantic features of all statements yields all categories. The category to which the target features of all target statements belong is obtained and denoted as the target category. The usage rate of the keyword is positively correlated with the number of target categories and the sum of the cluster densities of all target categories.
[0014] Preferably, the specific steps for constructing a keyword knowledge graph using the usage rate of keywords in the first response result from the joint recall results of multiple recalls are as follows:
[0015] For all keywords in the joint recall results, keywords with a usage rate greater than a preset usage rate threshold are selected. Each selected keyword is treated as a node, and any two nodes are connected. All nodes and the edges connecting the nodes constitute a keyword knowledge graph.
[0016] Preferably, the specific steps for obtaining the edge weights in the keyword knowledge graph are as follows:
[0017] For all statements in the first answer result, the statement in which the keywords of any two nodes in the keyword knowledge graph are located simultaneously is denoted as the first statement. The cosine similarity between the statement vector and the question vector corresponding to each first statement is denoted as the target feature of each first statement. The category to which the target features of all first statements belong is obtained and denoted as the first category. The edge weight between any two nodes is positively correlated with the number of first categories and the sum of the cluster densities of all first categories.
[0018] Preferably, the specific steps for supplementing the keyword knowledge graph by utilizing the usage rate of keywords in the second response result from each recall result are as follows:
[0019] Obtain the usage rate of each keyword in the second response result from each recall result; construct the first knowledge graph of each recall result using the usage rate of the keywords in the second response result; merge the first knowledge graph of each recall result into the keyword knowledge graph; then normalize the edge weight of each edge in the keyword knowledge graph to obtain the supplemented keyword knowledge graph.
[0020] Preferably, the specific steps for updating the recall metrics for each recall result based on the supplemented keyword knowledge graph are as follows:
[0021] The edge weight between any two nodes in the keyword knowledge graph is recorded as the association strength between the two nodes; for any text block in the common recall results, the nodes corresponding to all keywords in the keyword knowledge graph are obtained and recorded as the target nodes of each text block.
[0022] Obtain all nodes whose association strength with each target node is greater than a first preset threshold, and record them as associated nodes of each target node; use the usage rate of associated nodes of each target node and the usage rate of all target nodes to obtain the correction coefficient F1 for each text block; the formula for calculating the updated recall index of each text block is as follows:
[0023]
[0024] M1 represents the updated recall metric for each text block. This represents the recall metric for each text block before it is updated.
[0025] Preferably, the specific steps for obtaining the correction coefficient F1 for each text block using the usage rate of the associated nodes of each target node and the usage rate of all target nodes are as follows:
[0026] For each target node and all its associated nodes, the sum of the usage rates of all associated nodes is recorded as the information attention coefficient of each target node. Using the information attention coefficients of all target nodes, the usage rates of all target nodes are weighted and summed to obtain the correction coefficient for each text block.
[0027] Preferably, the specific steps for merging the first knowledge graph of each recall result into the keyword knowledge graph are as follows:
[0028] First, for nodes that are identical in both the first knowledge graph and the keyword knowledge graph, only the nodes in the keyword knowledge graph are retained. Nodes that exist only in the first knowledge graph but not in the keyword knowledge graph are added to the keyword knowledge graph. Then, for any pair of nodes in the first knowledge graph, the edge weight of the pair of nodes in the keyword knowledge graph is increased by an amount equal to the edge weight of the pair of nodes in the first knowledge graph. If the pair of nodes is not connected in the keyword knowledge graph, then the pair of nodes is connected, and the edge weight of the pair of nodes is set to the edge weight of the pair of nodes in the first knowledge graph.
[0029] Preferably, the specific steps following storing the user question and the supplemented keyword knowledge graph into the vector database are as follows:
[0030] The vector database stores question vectors of users' historical input questions, as well as keyword knowledge graphs corresponding to the question vectors;
[0031] When the user enters a question again, the question vector of the input question is recorded as the first vector; firstly, the keyword knowledge graph with the highest cosine similarity between the question vector and the first vector is retrieved from the vector database and recorded as the reference graph;
[0032] The steps before updating the recall index of each recall result based on the supplemented keyword knowledge graph are as follows: First, multiply the edge weights in the reference graph by the cosine similarity between the question vector corresponding to the reference graph and the first vector; then merge the reference graph into the supplemented keyword knowledge graph; and finally update the recall index of each recall result based on the keyword knowledge graph.
[0033] Another embodiment of the present invention provides an intelligent question-answering enhancement system based on RAG technology. The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor runs the computer program, it performs all the steps of the intelligent question-answering enhancement method based on RAG technology described above.
[0034] The beneficial effects of the technical solution of the present invention are:
[0035] This invention inputs the combined recall results from multiple paths into a large model to obtain a first answer result. A keyword knowledge graph is constructed using the usage rate of keywords from the combined recall results in the first answer result. Then, each path's recall result is input into the large model to obtain a second answer result. The keyword knowledge graph is supplemented based on the usage rate of keywords from each path's recall result in the second answer result. The usage rate describes how keywords from the combined recall results are used in the answer result, including whether the keyword is an important word in expressing semantics or describing logic, and whether it is a necessary or commonly used word in answering user questions. This process, through a multi-path strategy, ensures comprehensive coverage of potentially relevant text blocks from the vector database. Based on this, the usage rate of keywords is used for dynamic optimization driven by the keyword knowledge graph to deeply mine the logical and semantic relationships contained in the text blocks.
[0036] Furthermore, the recall metrics for each recall result are updated based on the supplemented keyword knowledge graph. Using the updated recall metrics, several text blocks are re-recalled from the common recall results of multiple recalls. These text blocks are then input into the large model to obtain the final answer. This process reconstructs and merges the common recall results of multiple recalls through the keyword knowledge graph, avoiding the following contradiction: the purpose of the recall process is to accurately retrieve text data for solving the problem from the plain text in the knowledge base. However, accurate retrieval of text data requires deep semantic and logical extraction of the question and plain text, which can only be achieved through the generation process. Finally, intelligent question answering is performed using recall methods with different advantages and disadvantages to ensure the completeness and reliability of the final answer. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 The flowchart illustrates the steps of an intelligent question-answering enhancement method based on RAG technology, as provided in an embodiment of the present invention. Detailed Implementation
[0039] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the intelligent question-answering enhancement method and system based on RAG technology proposed in this invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0041] The following description, in conjunction with the accompanying drawings, details the specific scheme of the intelligent question-answering enhancement method and system based on RAG technology provided by this invention.
[0042] Example 1:
[0043] Please see Figure 1 The diagram illustrates a flowchart of a method for enhancing intelligent question answering based on RAG technology according to an embodiment of the present invention. The method includes the following steps:
[0044] Step S101: Construct a vector database for the knowledge base, and perform multi-way retrieval from the vector database based on user questions.
[0045] In this embodiment, the knowledge base is a Word document. The plain text content, including author, creation date, chapter titles, headings at each level, and main content, is read from the file. In other embodiments, the knowledge base can also be a PDF, Excel, Markdown, HTML, or other format files, from which the plain text content is read or parsed.
[0046] Furthermore, the plain text content of each file in all files within the knowledge base is divided into blocks to obtain several text blocks.
[0047] As an example, the plain text content of each file is divided into blocks to obtain several text blocks. The methods include:
[0048] Divide the text into blocks according to punctuation marks, that is, use periods, semicolons, etc. as dividing points to divide the plain text content into several text blocks.
[0049] In some embodiments, the text can be divided into blocks according to paragraphs or chapters. The methods for text segmentation are existing technologies, and this embodiment will not elaborate on or limit them.
[0050] In this embodiment, the maximum number of tokens for each text block is set to N. This embodiment uses N=512 as an example for description. In other embodiments, N can be set to other integer values, and its preferred value range is [256, 2048].
[0051] Specifically, if a text block contains more than N tokens, then the text block is divided into two text blocks, including:
[0052] Divide the text block into several sub-segments (e.g., using commas to divide the text block into several sub-segments), and denote the number of sub-segments as n0. Let... , The symbol indicates rounding up. In this embodiment, the first n+2 sub-segments are treated as one text block, and the last n-2 sub-segments are treated as another text block. These two text blocks have 4 overlapping sub-segments. In other embodiments, the integer 2 in this process can be replaced with other integers greater than or equal to 1 to ensure that the two text blocks have appropriate overlap.
[0053] Furthermore, Sentence-BERT technology is used to encode (or embed) each text block into a vector, denoted as a text block vector. Each text block in the knowledge base and its corresponding text block vector are stored in a vector database. In this implementation, the dimension of each text block vector is set to 128 dimensions.
[0054] Next, the system reads the text-formatted question input by the user and encodes it as a vector, denoted as the question vector. It should be noted that if the text-formatted question input by the user exceeds N tokens, the excess tokens are deleted, and the user is prompted that "input is too long."
[0055] The problem is used to retrieve data from a vector database. In this embodiment, a multi-path retrieval method is used, which includes keyword retrieval, semantic similarity retrieval, and text summary retrieval.
[0056] Each recall result contains several text blocks, and each text block corresponds to a recall metric.
[0057] Specifically, as an example, the keyword recall method is as follows:
[0058] The question is segmented into words to obtain several keywords. Each text block in the vector database is treated as a file. The TF-IDF value of each keyword in each file is calculated based on all files. The average TF-IDF value of all keywords in all files is taken as the recall index for each text block. The higher the recall index, the more usable the text block is for answering the user's input question, and the more it needs to be recalled (i.e., the more it needs to be retrieved from the vector database).
[0059] This example recalls the top K1 text blocks with the largest recall metrics, that is, retrieves the top K1 text blocks with the largest recall metrics as the recall results.
[0060] Among them, calculating the TF-IDF value of each keyword in each file is a well-known technology, and this embodiment will not elaborate specifically. The word segmentation in this embodiment uses the jieba word segmentation technology. It should be noted that when using jieba word segmentation, a stop word library needs to be set manually. The stop word library contains words with relatively low contribution to text understanding, such as "among them", "currently", etc.
[0061] The advantages of keyword recall are: it can accurately hit the text information describing the problem. The disadvantages are: text blocks that can solve the problem but do not contain keywords cannot be recalled.
[0062] As an example, the method of semantic similarity recall is:
[0063] Obtain the cosine similarity between the problem vector and the text block vectors of each text block in the vector database, which is also recorded as the recall metric of each text block. The larger the recall metric, the closer the semantics of the text block is to the semantics of the user input problem, and the more it needs to be recalled.
[0064] This example recalls the top K2 text blocks with the largest recall metrics, that is, retrieves the top K2 text blocks with the largest recall metrics as the recall results.
[0065] The advantages of semantic similarity recall are: it can describe whether the text information fits the problem from the semantic level. The disadvantages are: the text block vectors are not sufficient to describe all aspects of the semantics of the text blocks, especially when the semantic information of the text blocks is complex, resulting in the inability to accurately recall all the text blocks of the user input problem.
[0066] As an example, the method of text summary recall is:
[0067] Obtain the summary of each text block. In this embodiment, the TextRank algorithm is used to obtain the text summary. In other embodiments, the TF-IDF algorithm can be used to obtain the summary. In some embodiments, large models such as GPT4 can also be directly used to obtain the summary.
[0068] Encode the summary into a vector, denoted as the summary vector. The cosine similarity between the summary vector and the problem vector is denoted as the recall metric of each text block. This example recalls the top K3 text blocks with the largest recall metrics, that is, retrieves the top K3 text blocks with the largest recall metrics as the recall results.
[0069] The advantage of text summarization recall lies in leveraging the semantics of key information within text blocks. When dealing with text blocks that are information-rich and complex, it can avoid interference from unimportant information. However, its drawback is that it discards some information from the text block, potentially omitting content relevant to describing the problem.
[0070] In this embodiment, topK1, topK2, and topK3 are each set to 10.
[0071] In other embodiments, multi-path recall also includes semantic similarity recall at different contextual scales, specifically as follows:
[0072] Multi-scale segmentation of plain text content, including:
[0073] Divide the text into blocks according to punctuation marks, that is, use periods, semicolons, etc. as dividing points to divide the plain text content into several sentences.
[0074] Each pair of adjacent statements is treated as a text block, and all resulting text blocks are denoted as the chunking result under the first context scale. Each group of three adjacent statements is treated as a text block, and all resulting text blocks are denoted as the chunking result under the second context scale. These text blocks have a larger context scale, and duplicate content is allowed between different text blocks.
[0075] The text blocks in the results for each context scale are encoded and stored separately to obtain a vector database for each context scale.
[0076] Semantic similarity recall is performed using vector databases at each context scale to obtain recall results at each context scale. The process of obtaining recall results at each context scale is considered as one of the recall processes in the multi-path recall.
[0077] The advantage of semantic similarity recall at different contextual scales is that it can capture textual information that is semantically related but distributed in different text locations. The disadvantage is that the contextual scale is large, the semantic information is abundant and complex, and it is easy to overlook some details.
[0078] In summary, each of the above recall methods is a well-known method. Different recall methods have different advantages and disadvantages. The recall results obtained by multi-way recall (that is, using multiple recall methods at the same time) are called common recall results. These common recall results can be used to obtain more accurate and reliable response results.
[0079] Step S102: Input the combined recall results from multiple recalls into the large model to obtain the first answer result. Construct a keyword knowledge graph using the usage rate of keywords in the first answer result from the combined recall results of multiple recalls.
[0080] The common recall result of the multi-path recall refers to all text blocks in the recall results obtained by all recall methods.
[0081] The user-input question and the co-recall results are input together into the large model, which outputs the first answer. Note that the co-recall results may contain identical text blocks or content, even though these text blocks or content come from different recall methods. When inputting the co-recall results into the large model, duplicate text content needs to be removed before inputting into the large model. In this embodiment, the large model uses the GPT4 model; other embodiments may also use the DeepSeek model, and this embodiment does not impose a specific limitation.
[0082] It should be noted that the process of obtaining the recall results described above is called the recall process, and the process of obtaining the first answer result based on the recall results is called the generation process. The processing of text information (i.e., text blocks and user-input questions) during the recall process is only shallow feature processing; for example, encoding into vectors and similarity calculations are both shallow feature processing. The generation process, on the other hand, involves deep feature extraction of text information (i.e., text blocks in the recall results and user-input questions). The former cannot accurately uncover the precise correlation and logical relationship between the user-input questions and text blocks, while the latter, based on a large model, can deeply mine the logical and semantic relationships contained in the text information. For example, it can uncover the deeper meaning and true intent of the user-input questions, extract which semantic information in the text blocks can describe or solve the problem, or is related to the problem, and split and reassemble text information from different text blocks or different positions within the same text block to adapt to the semantics expressed by the user-input questions. In summary, the purpose of the recall process is to accurately retrieve textual data for solving problems from the plain text in the knowledge base. However, accurate retrieval of textual data requires deep semantic and logical extraction of the problem and the plain text, which can only be achieved through the generation process. Based on this, this embodiment reconstructs the recall process based on the first answer result.
[0083] On the other hand, it should be noted that while the joint recall results already include text blocks that can solve the problem, they also contain a large amount of irrelevant, interfering, and useless information. This information can cause the first response to be incomplete, redundant, or incomplete, making it unsuitable as the final response result of RAG technology. However, the logical and semantic relationships between keywords in the first response result can reflect and adapt to the user's question.
[0084] In this embodiment, in order to reconstruct the recall process based on the first answer result, a keyword knowledge graph is first constructed using the usage rate of keywords in the first answer result from the common recall results of multiple recalls.
[0085] The usage rate of keywords in the first response result from the combined recall results describes how often the keywords in the combined recall results are used in the first response result. This includes whether the keywords are important words used in the first response result to express semantics or describe logic, or whether they are necessary or commonly used words in the first response result when answering the user's question. The higher the usage rate, the more likely the keywords are to be important words used in the first response result to express semantics or describe logic, or necessary or commonly used words in the first response result when answering the user's question.
[0086] The keyword knowledge graph is used to describe the relationships between keywords with different usage rates. In this embodiment, the recall process will be reconstructed using the keyword knowledge graph.
[0087] As an alternative example, methods for obtaining the usage rate of keywords in the first response result in the co-recall results of multi-path recall include:
[0088] For any keyword in the common recall results of multi-path recall, the common recall results of multi-path recall are treated as one file, and each text block outside the common recall results in the vector database is also treated as a separate file. The TF-IDF value of the keyword in the common recall results is obtained using all files, and the frequency of the keyword in the first response result is obtained. The product of the TF-IDF value and the frequency is recorded as the usage rate.
[0089] As a preferred example, methods for obtaining the usage rate of keywords in the first response result in the joint recall results of multi-path recall include:
[0090] For any keyword in the common recall results of multi-path recall, and for any sentence in the first response result (identified by periods and semicolons), each sentence in the first response result is encoded into a vector, denoted as the sentence vector, and the cosine similarity between the sentence vector and the question vector is calculated, denoted as the semantic feature of each sentence. The semantic features of all sentences are obtained, and the distribution of these semantic features is used to describe the semantic or logical distribution of the first response result.
[0091] Statements containing the keyword are categorized as target statements, and the cosine similarity between the statement vector and the question vector corresponding to each target statement is categorized as the target feature of each target statement.
[0092] Perform K-Means clustering on the semantic features of all statements, and set the number of the resulting clusters as K1. Obtain the category to which the target features of all target statements belong, and denote it as the target category. Obtain the ratio of the number of target categories to K1, and denote it as the first ratio. Obtain the cluster density of each target category. The product of the sum of the cluster densities of all target categories and the first ratio is denoteed as the usage rate of the keyword.
[0093] The larger the sum of cluster densities and the first ratio, the more important the target sentence containing the keyword is in the first answer result, and the more important it is in the semantic or logical expression of the first answer result. In this case, the usage rate is relatively high.
[0094] In this embodiment, K1 is set to one-fifth of the number of statements in the first answer result (rounded up). The method for obtaining the cluster density of each target category in this embodiment is as follows: the standard deviation of all semantic features within each target category is denoted as x, and exp(-x) is used as the cluster density, where exp() represents an exponential function with the natural constant as the base.
[0095] As another example, methods for obtaining the usage rate of keywords in the first response result in the joint recall results of multi-path recall include:
[0096] The product of the usage rates obtained from the optional and preferred examples above is used as the usage rate obtained in this example.
[0097] Specifically, if the keyword does not exist in the first answer result, then the keyword usage rate is set to 0.
[0098] At this point, the usage rate of each keyword in the joint recall results was obtained.
[0099] As an optional example, a keyword knowledge graph can be constructed using the usage rate of keywords in the first response result from the common recall results of multi-path recall. The methods include:
[0100] For all keywords in the combined recall results of multiple recalls, keywords with a usage rate greater than a preset usage rate threshold are selected. This embodiment uses a preset usage rate threshold of 0.05 as an example. In other embodiments, the preset usage rate threshold can be set to 0, and the preferred value range is [0, 0.5].
[0101] Each selected keyword is treated as a node, and any two nodes are connected by an edge. All nodes and the edges connecting them constitute a keyword knowledge graph, with each edge having a weight of 1.
[0102] As a preferred example, a keyword knowledge graph is constructed using the usage rate of keywords in the first response result from the common recall results of multi-path recall. The methods include:
[0103] Based on the keyword knowledge graph obtained from the above optional examples, for all statements in the first answer result, when the keywords corresponding to any two nodes in the keyword knowledge graph are simultaneously in the same statement, the statement is recorded as the first statement. At this time, the keywords corresponding to any two nodes are simultaneously in one or more first statements. The cosine similarity between the statement vector and the question vector corresponding to each first statement is recorded as the target feature of each first statement.
[0104] Obtain the category to which all target features of the first statement belong, denoted as the first category. Obtain the ratio of the number of first categories to K1, denoted as the second ratio. Obtain the cluster density of each first category. The product of the sum of the cluster densities of all first categories and the second ratio is denoted as the bonding strength between any two nodes.
[0105] The stronger the combination, the more important the sentence containing the keywords of any two nodes is to the semantic or logical expression of the first answer result, and the more the keywords of the two nodes can be combined to answer the user's question.
[0106] If the keywords corresponding to any two nodes are not in the same statement at the same time, then the combination strength of any two nodes is set to 0.
[0107] Obtain the edge weights of any two nodes, sum these edge weights with the associative strength, and use the sum as the edge weights of any two nodes again.
[0108] Step S103: Input each recall result into the large model to obtain the second answer result, and use the usage rate of keywords in each recall result in the second answer result to supplement the keyword knowledge graph.
[0109] The recall results for each recall (referred to as each recall result) and the user's input question are input into the large model to obtain the second response result for each recall.
[0110] It should be noted that the aforementioned joint recall results contain a large amount of irrelevant, interfering, and useless information. This information can lead to the first response being incomplete, redundant, or missing. In contrast, the recall results from each individual recall are concise, and important information is less likely to be buried by redundant information. Therefore, the second response may contain important words that are not present in the first response. Thus, this embodiment uses the usage rate of keywords from each recall result in the second response to supplement the keyword knowledge graph. Furthermore, it should be noted that since the recall results from each individual recall may not contain all the information needed to answer the user's question, the second response cannot be directly used as the final response result of RAG technology.
[0111] As an example, the keyword knowledge graph is supplemented based on the usage rate of keywords in the second response result from each recall result. This includes methods such as:
[0112] First, obtain the usage rate of each keyword in the second response result from each recall result.
[0113] Then, the first knowledge graph of each recall result is constructed using the usage rate of keywords in the second response result (similar to step S102).
[0114] The first knowledge graph is merged into the keyword knowledge graph. The merging method is as follows: First, for nodes that are the same in both the first and keyword knowledge graphs, only the nodes in the keyword knowledge graph are retained. Then, for nodes that exist only in the first knowledge graph and not in the keyword knowledge graph, they are added to the keyword knowledge graph. Next, for any pair of nodes in the first knowledge graph, the pair of nodes is found in the keyword knowledge graph after the nodes are added, and the edge weight of the pair of nodes in the keyword knowledge graph is increased by an amount equal to the edge weight of the pair of nodes in the first knowledge graph. If the pair of nodes is not connected in the keyword knowledge graph, then the pair of nodes is connected, and the edge weight of the pair of nodes is set to the edge weight of the pair of nodes in the first knowledge graph.
[0115] Using the above method, the first knowledge graphs of all path recall results are obtained, and all the first knowledge graphs of all path recall results are merged into the keyword knowledge graph.
[0116] In the merged keyword knowledge graph, if there is any node that is not connected to any other node, then connect that node to all nodes respectively, and set the edge weight of the connected edge to 1.
[0117] The edge weights of each edge in the keyword knowledge graph are normalized; this embodiment uses the softmax formula for normalization. The edge weights between any two nodes in the keyword knowledge graph are denoted as the association strength between those two nodes.
[0118] Step S104: Update the recall metrics for each recall result based on the supplemented keyword knowledge graph.
[0119] For any text block in the common recall results, obtain all keywords in that text block, and obtain the nodes corresponding to these keywords in the keyword knowledge graph, which are recorded as the target nodes of each text block.
[0120] All nodes whose association strength with each target node is greater than a first preset threshold th1 are identified and recorded as associated nodes of each target node. The sum of the usage rates of all associated nodes is recorded as the information attention coefficient of each target node. The larger the information attention coefficient, the more effectively the keywords corresponding to the target node can solve or describe the user's question when combined with other keywords. In this embodiment, th1=0.1; in other embodiments, it can be set to other values, preferably within the range of (0, 1).
[0121] By using the information attention coefficients of all target nodes, the usage rates of all target nodes are weighted and summed to obtain the correction coefficient for each text block. The larger the correction coefficient, the more important the text information in each text block is to the semantic or logical expression of the first answer result. It can be combined with other text information to answer the user's question, or in other words, each text block contains the semantics and logic that are indispensable for answering the user's question.
[0122] The updated recall metric for each text block is positively correlated with the correction coefficient for each text block.
[0123] As an example, by using the information attention coefficients of all target nodes, the usage rates of all target nodes are weighted and summed to obtain the correction coefficient for each text block. The formula includes:
[0124]
[0125] in This represents the information attention coefficient of the j-th target node in each text block. Let Q represent the utilization rate of the j-th target node in each text block, and let Q represent the number of target nodes in each text block. This represents the correction factor for each text block.
[0126] Specifically, if each text block does not have a corresponding target node, then F1 is set to 0.
[0127] As an example, the formula for calculating the updated recall metric for each text block is:
[0128]
[0129] M1 represents the updated recall metric for each text block. This represents the recall index before the update of each text block (or the recall index obtained in step S101).
[0130] Step S105: Use the updated recall metrics to recall several text blocks from the combined recall results of multiple recalls, and input these text blocks into the large model to obtain the final answer result.
[0131] For all text blocks in the common recall results of multi-path recall, the top K text blocks with the largest updated recall metrics are obtained and denoted as the recall text blocks. The recall text blocks and the user's question are input into the large model, the large model outputs the final answer result, and the final answer result is displayed to the user.
[0132] In other embodiments, after obtaining the recall text blocks, a cross-encoder can be used to reorder the recall text blocks. The top N reordered recall text blocks and the user's question are then input into a large model. The large model outputs the final answer, which is then displayed to the user. In this embodiment, topK=10; in other embodiments, it can be set to other values, preferably within the range of [5, 50]. The reordering process described above is a well-known technique and will not be elaborated upon in this embodiment.
[0133] It should be noted that the common recall results contain the same text blocks. These text blocks come from different recall methods (the recall metrics obtained in step S101 may also be different). The updated recall metrics may be different. If there are identical text blocks among the top K text blocks with the largest recall metrics, the duplicate text blocks (or each text block contained in other text blocks) will be deleted. Assuming K2 text blocks are deleted, then the K2 text blocks with the largest recall metrics will be selected again from all the remaining text blocks so that the recalled text blocks make up the top K.
[0134] This concludes the implementation. This implementation inputs the combined recall results from multiple paths into a large model to obtain the first answer result. A keyword knowledge graph is constructed using the usage rate of keywords from the combined recall results in the first answer result. Then, each path's recall result is input into the large model to obtain the second answer result. The keyword knowledge graph is supplemented based on the usage rate of keywords from each path's recall result in the second answer result. The usage rate describes how keywords from the combined recall results are used in the answer result, including whether the keyword is an important word in expressing semantics or describing logic, and whether it is a necessary or commonly used word in answering user questions. This process, through a multi-path strategy, ensures comprehensive coverage of potentially relevant text blocks from the vector database. Based on this, the usage rate of keywords is used to deeply mine the logical and semantic relationships contained within the text blocks through dynamic optimization driven by the keyword knowledge graph.
[0135] Furthermore, the recall metrics for each recall result are updated based on the supplemented keyword knowledge graph. Using the updated recall metrics, several text blocks are re-recalled from the common recall results of multiple recalls. These text blocks are then input into the large model to obtain the final answer. This process reconstructs and merges the common recall results of multiple recalls through the keyword knowledge graph, avoiding the following contradiction: the purpose of the recall process is to accurately retrieve text data for solving the problem from the plain text in the knowledge base. However, accurate retrieval of text data requires deep semantic and logical extraction of the question and plain text, which can only be achieved through the generation process. Finally, intelligent question answering is performed using recall methods with different advantages and disadvantages to ensure the completeness and reliability of the final answer.
[0136] Example 2:
[0137] The difference between this embodiment and Embodiment 1 is as follows:
[0138] After the large model outputs the final answer, step S106 is included, specifically:
[0139] Obtain the usage rate of keywords in all recalled text blocks in the final answer results; supplement the keyword knowledge graph based on the usage rate of keywords in all recalled text blocks in the final answer results (the specific process is the same as S103).
[0140] The user-input questions, input question vectors, and supplemented keyword knowledge graphs are stored in a vector database.
[0141] This concludes step S106.
[0142] When the user enters a question again, the question vector of the entered question is recorded as the first vector.
[0143] In the vector database, the question vector with the highest cosine similarity to the first vector is obtained. If the question vector is less than 0.8, the final answer result is obtained using the method in Example 1.
[0144] When the question vector is greater than or equal to 0.8, the keyword knowledge graph for that question vector is retrieved from the vector database and recorded as the reference graph. Then, the keyword knowledge graph is re-acquired according to steps S102 to S103 of Embodiment 1, and the reference graph is merged into the keyword knowledge graph (see step S103 for details). Then, the final answer result is obtained according to steps S104 to S105. Finally, step S106 of this embodiment is re-executed.
[0145] In other embodiments, the edge weights of all edges in the reference graph are multiplied by the cosine similarity between the question vector corresponding to the reference graph and the first vector, and then the reference graph is merged into the keyword knowledge graph. This embodiment is used to avoid excessive interference of the reference graph on the keyword knowledge graph, especially to avoid the problem that the reference graph has low reference value and introduces irrelevant redundant information into the keyword knowledge graph when the similarity between the question corresponding to the reference graph and the input question is small.
[0146] In some embodiments, when the question vector is greater than or equal to 0.95, the reference graph is directly used as the keyword knowledge graph in step S104, and the final answer result is obtained using the methods in steps S104 to S105.
[0147] In the above process, as users continue to ask questions, the keyword knowledge graph in the vector database will be continuously updated. The keyword knowledge graph in the vector database corresponds to the question vector of the historical questions. The keyword knowledge graph is used to fuse the multi-way recall results to obtain the final answer, making the final answer more reliable and accurate.
[0148] Specifically, when the number of keyword knowledge graphs stored in the vector database exceeds 1000, the first 10 keyword knowledge graphs stored in the vector database will be deleted to avoid the data volume in the vector database becoming too large.
[0149] Example 3:
[0150] This embodiment provides an intelligent question-answering enhancement system based on RAG technology. The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor runs the computer program, it performs all the steps in all the embodiments described above.
[0151] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An intelligent question answering enhancement method based on RAG technology, characterized in that, The method comprises the following steps: The vector database of the knowledge base is constructed, and multi-path recall is performed from the vector database according to the user question, wherein the multi-path recall at least comprises keyword recall, semantic similarity recall and text summary recall; each path recall result comprises a plurality of text blocks retrieved from the vector database, and each text block in each path recall result corresponds to a recall index; The common recall result of the multi-path recall is input into the large model to obtain a first answer result, and a keyword knowledge graph is constructed by using the usage rate of the keywords in the common recall result of the multi-path recall in the first answer result; each path recall result is input into the large model to obtain a second answer result, and the keyword knowledge graph is supplemented by using the usage rate of the keywords in the second answer result; The recall index of each path recall result is updated according to the supplemented keyword knowledge graph, a plurality of text blocks are re-recalled from the common recall result of the multi-path recall by using the updated recall index, and the plurality of text blocks are input into the large model to obtain a final answer result; The keyword knowledge graph is supplemented again by using the usage rate of the keywords in the plurality of text blocks in the final answer result, and the user question and the supplemented keyword knowledge graph are stored into the vector database; The specific steps of constructing the keyword knowledge graph by using the usage rate of the keywords in the common recall result of the multi-path recall in the first answer result comprise the following steps: For all keywords in the common recall result, keywords with a usage rate greater than a preset usage rate threshold are screened out, each screened keyword is taken as a node, any two nodes are connected, and all nodes and edges connecting the nodes constitute the keyword knowledge graph; The specific steps of supplementing the keyword knowledge graph by using the usage rate of the keywords in the second answer result in each path recall result comprise the following steps: The usage rate of each keyword in the second answer result in each path recall result is obtained; a first knowledge graph of each path recall result is constructed by using the usage rate of the keywords in the second answer result in each path recall result; the first knowledge graph of each path recall result is merged into the keyword knowledge graph respectively; and then the edge weight of each edge in the keyword knowledge graph is normalized to obtain the supplemented keyword knowledge graph.
2. The method of claim 1, wherein the RAG-based intelligent question answering enhancement method is characterized by, The specific steps of obtaining the usage rate of the keywords in the common recall result of the multi-path recall in the first answer result comprise the following steps: For any keyword in the common recall result and for all sentences in the first answer result; each sentence in the first answer result is coded into a sentence vector, and the user question is coded into a question vector, and the cosine similarity of the sentence vector and the question vector is calculated, which is recorded as the semantic feature of each sentence; Each sentence containing the keyword is recorded as a target sentence, and the cosine similarity of the sentence vector and the question vector corresponding to each target sentence is recorded as the target feature of each target sentence; The semantic features of all the sentences are clustered to obtain all categories, and the category to which the target feature of all target sentences belongs is obtained and recorded as a target category, and the use rate of the keyword is positively correlated with the number of target categories and the sum of the clustering densities of all target categories.
3. The method of claim 1, wherein the RAG-based intelligent question answering enhancement method is characterized by, The specific steps for obtaining the edge weight value in the keyword knowledge graph are as follows: For all the sentences in the first answer result, the sentences in which the keywords corresponding to any two nodes in the keyword knowledge graph are located are recorded as first sentences, and the cosine similarity between the sentence vector and the question vector corresponding to each first sentence is recorded as the target feature of each first sentence; The categories to which the target features of all the first sentences belong are obtained and recorded as first categories, and the edge weight value between any two nodes is positively correlated with the number of first categories and the sum of the clustering densities of all first categories.
4. The method of claim 1, wherein the RAG-based intelligent question answering enhancement method is characterized by, The specific steps for updating the recall index of each recall result according to the supplemented keyword knowledge graph are as follows: The edge weight value between any two nodes in the keyword knowledge graph is recorded as the association strength between the two nodes; for any one text block in the common recall result, the nodes corresponding to all the keywords in the text block in the keyword knowledge graph are obtained and recorded as target nodes of each text block; All the nodes associated with each target node and having an association strength greater than a first preset threshold are obtained and recorded as associated nodes of each target node; the use rate of each target node and the use rate of all target nodes are used to obtain a correction coefficient F1 of each text block; and the calculation formula of the updated recall index of each text block is: M1 represents the updated recall indicator for each text block, M0 represents the pre-update recall indicator for each text block.
5. The method of claim 4, wherein the RAG-based intelligent question answering enhancement method is characterized by, The specific steps for obtaining the correction coefficient F1 of each text block by using the use rate of each target node and the use rate of all target nodes are as follows: For all the associated nodes of each target node, the sum of the use rates of all the associated nodes is recorded as an information attention coefficient of each target node; the use rates of all the target nodes are weighted and summed by using the information attention coefficients of all the target nodes to obtain the correction coefficient of each text block.
6. The method of claim 1, wherein the RAG-based intelligent question answering enhancement method is characterized by, The specific steps for merging the first knowledge graph of each recall result into the keyword knowledge graph are as follows: First, for the same nodes in the first knowledge graph and the keyword knowledge graph, only the nodes in the keyword knowledge graph are retained; the nodes that exist only in the first knowledge graph and do not exist in the keyword knowledge graph are added to the keyword knowledge graph; then for any pair of nodes in the first knowledge graph; the edge weight value of the pair of nodes in the keyword knowledge graph is increased, and the increase amount is equal to the edge weight value of the pair of nodes in the first knowledge graph; if the pair of nodes is not connected in the keyword knowledge graph, the pair of nodes is connected, and the edge weight value of the pair of nodes is set to the edge weight value of the pair of nodes in the first knowledge graph.
7. The method of claim 1, wherein the RAG-based intelligent question answering enhancement method is characterized by, The specific steps for storing the user question and the supplemented keyword knowledge graph into the vector database are as follows: The vector database stores the question vector of the user's historical input question and the keyword knowledge graph corresponding to the question vector. When the user inputs the question again, the question vector of the input question is denoted as a first vector; first, the keyword knowledge graph with the maximum cosine similarity between the question vector and the first vector is retrieved from the vector database, and is denoted as a reference graph; The steps included before the step of updating the recall indexes of each of the recall results according to the supplemented keyword knowledge graph include: first, multiplying the edge weight values in the reference graph by the cosine similarity between the question vector corresponding to the reference graph and the first vector; then, merging the reference graph into the supplemented keyword knowledge graph; and finally, updating the recall indexes of each of the recall results according to the keyword knowledge graph.
8. An intelligent question answering enhancement system based on RAG technology, the system comprising: Memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes all steps of the intelligent question answering enhancement method based on the RAG technology according to any one of claims 1-7 when executing the computer program.
Citation Information
Patent Citations
Intelligent question answering method, device and equipment based on large language model and storage medium
CN119357359A
Large language model knowledge base question answering system based on multi-path fusion recall retrieval algorithm
CN120144773A
Cited By
Intelligent question-answering system enhancement method based on retrieval enhancement generation technology
CN122287859A