Multi-round network retrieval self-error correction method, intelligent question and answer method based on multi-round network retrieval self-error correction and related devices
By employing a multi-round network retrieval self-correction method and sentence-level RAG knowledge base optimization, the problem of incorrect answers caused by the closed nature of the knowledge base in intelligent question answering systems was solved, achieving efficient and accurate knowledge block retrieval and improved question answering accuracy.
Patent Information
- Application Number
- CN202510974957.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-31
AI Technical Summary
In existing intelligent question answering systems, the closed nature of the local RAG knowledge base may lead to biased knowledge blocks retrieved, resulting in incorrect answers. Furthermore, existing methods are time-consuming and inefficient, failing to efficiently and accurately answer users' time-sensitive questions.
A multi-round network retrieval self-correction method is adopted, which performs different rounds of network retrieval and correction processes through similarity matching and preset triggering mechanisms. Combined with sentence-level RAG knowledge base optimization and similarity tag generation model, the accuracy and efficiency of knowledge blocks are improved.
It improves the relevance and accuracy between knowledge blocks and questions, enhances the question-answering accuracy of the large language model, reduces the possibility of incorrect answers, and improves the system's response efficiency.
Smart Images

Figure CN120873285A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically a multi-round network retrieval self-correction method, an intelligent question-answering method based on multi-round network retrieval self-correction, and related devices. Background Technology
[0002] In the digital age, intelligent question-answering systems have become core tools for information retrieval and intelligent question answering in cross-industry applications. With the advent of large language models, their powerful instruction understanding and text generation capabilities have been impressive, leading to superior question-answering performance in intelligent question-answering assistants built upon them. However, due to the vast amounts of data used in their pre-training, large models inevitably suffer from "illusions" when faced with time-sensitive questions, ultimately resulting in low accuracy. Therefore, in the context of using large language models for intelligent knowledge question answering, efficiently and accurately answering users' time-sensitive questions has become a critical challenge that urgently needs to be addressed.
[0003] Currently, many technologies utilize Retrieval Augmentation (RAG) based on large language models. This framework pre-injects professional knowledge into a local RAG knowledge base to guide the large language model in answering questions. Specifically, it retrieves relevant knowledge blocks from the local RAG knowledge base based on the question and then inputs these knowledge blocks into the large language model to answer the question. However, due to the closed nature of the local RAG knowledge base, the retrieved knowledge blocks may be biased, leading to incorrect answers that mislead users. Therefore, users need to continuously inject professional knowledge into the knowledge base to support the large language model in accurately answering questions. However, this process is time-consuming, inefficient, and has a certain degree of latency. Summary of the Invention
[0004] This invention provides a multi-round network retrieval self-correction method, an intelligent question answering method based on multi-round network retrieval self-correction, and related devices, which overcomes the shortcomings of the prior art. It can effectively solve the problem that the closed nature of the local RAG knowledge base may lead to deviations in the retrieved knowledge blocks, resulting in incorrect answers that mislead users' questions.
[0005] One of the technical solutions of this invention is achieved through the following measures: a multi-round network retrieval self-correction method, comprising:
[0006] A similarity match is performed between a given question and the retrieved preliminary answers to obtain a similarity level. The preliminary answers include one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and no similarity.
[0007] In response to high similarity, it does not trigger a multi-round network retrieval self-correction process;
[0008] In response to similarity levels other than high similarity, different rounds of network retrieval self-correction processes are triggered according to preset triggering mechanisms, including:
[0009] If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results.
[0010] If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time.
[0011] If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
[0012] The following are further optimizations and / or improvements to the above-mentioned technical solution:
[0013] The above-mentioned similarity matching between a given question and the retrieved answers is accomplished using a similarity label generation model. The construction process of the similarity label generation model includes:
[0014] A fine-tuning dataset was generated using the ChatGPT 3.5 model. The fine-tuning dataset includes several modeling datasets, each of which includes a question, the knowledge blocks that the question may retrieve, and the similarity between the question and the knowledge blocks.
[0015] Supervised fine-tuning training of the pre-trained model is performed using the fine-tuning dataset to obtain a similarity label generation model, where the pre-trained model is a small model obtained by training using the training dataset.
[0016] The above also includes obtaining a given question and the preliminary answers retrieved, including:
[0017] Obtain the questions submitted by users and use them as given questions;
[0018] Retrieve one or more knowledge blocks most relevant to the question from the sentence-level RAG knowledge base to form a preliminary answer.
[0019] The optimization strategies for the above sentence-level RAG knowledge base include:
[0020] The documents to be processed are divided into sentence-level segments;
[0021] The sliding window method is used to determine the merging of two adjacent sentences by combining analytical indicators, including semantic coherence, semantic integrity, and semantic clarity.
[0022] The second technical solution of the present invention is achieved through the following measures: an intelligent question-answering method based on multi-round network retrieval self-correction, comprising:
[0023] Given a problem, obtain knowledge blocks for the given problem through a multi-round network retrieval self-correction method;
[0024] Input a given question and knowledge block into the question answering model to get the final answer to the given question. The question answering model is based on a large language model.
[0025] The following are further optimizations and / or improvements to the above-mentioned technical solution:
[0026] The above question-answering models include:
[0027] Word embedding models vectorize the given input question and knowledge block;
[0028] The reordering model reorders knowledge blocks and uses the Sort function to score the similarity between the knowledge blocks and the given question, selecting the Top_K knowledge blocks with the highest similarity.
[0029] The large language model analyzes a given question, integrates knowledge blocks, and generates the final answer.
[0030] The third technical solution of the present invention is achieved through the following measures: a multi-round network retrieval self-correction device, comprising:
[0031] The similarity level determination unit performs similarity matching between a given question and the retrieved preliminary answer to obtain a similarity level. The preliminary answer includes one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and no similarity.
[0032] Network retrieval routing unit, including:
[0033] In response to high similarity, it does not trigger a multi-round network retrieval self-correction process;
[0034] In response to similarity levels other than high similarity, different rounds of network retrieval self-correction processes are triggered according to preset triggering mechanisms, including:
[0035] If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results.
[0036] If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time.
[0037] If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
[0038] The fourth technical solution of the present invention is achieved through the following measures: an intelligent question-answering device based on multi-round network retrieval self-correction, comprising:
[0039] The basic data acquisition unit acquires a given question and obtains knowledge blocks for the given question through a multi-round network retrieval self-correction method;
[0040] The question-answering analysis unit takes a given question and knowledge block as input to the question-answering model and obtains the final answer to the given question. The question-answering model is based on a large language model.
[0041] The fifth technical solution of the present invention is achieved through the following measures: an electronic device, including a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the steps in the intelligent question answering method based on multi-round network retrieval self-correction.
[0042] The sixth technical solution of the present invention is achieved through the following measures: a storage medium, characterized in that the storage medium stores a computer program that can be read by a computer, the computer program being configured to execute the steps in the intelligent question answering method based on multi-round network retrieval self-correction when running.
[0043] This invention utilizes a fine-tuning dataset to supervise the fine-tuning of a pre-trained model to obtain a similarity label generation model. Based on this model, the similarity matching result between a given question and the retrieved preliminary answer is determined, triggering different rounds of network retrieval self-correction. This significantly improves the accuracy and efficiency of acquiring relevant knowledge blocks through network retrieval. Furthermore, multi-round network retrieval self-correction enhances the relevance between the retrieved knowledge blocks and the question, facilitating improved question-answering accuracy in subsequent large language models. Simultaneously, this invention employs a sentence-level RAG knowledge base optimization strategy to reconstruct the knowledge base architecture, making the preliminary answers retrieved from the knowledge base more accurate. Attached Figure Description
[0044] Appendix Figure 1 This is a schematic diagram of a multi-round network retrieval self-correction method provided by the present invention.
[0045] Appendix Figure 2 This is a schematic diagram of another multi-round network retrieval self-correction method provided by the present invention.
[0046] Appendix Figure 3 A schematic diagram illustrating the sentence-level RAG knowledge base optimization strategy provided by this invention.
[0047] Appendix Figure 4 This is a schematic diagram of the intelligent question-answering method provided by the present invention.
[0048] Appendix Figure 5 This is a schematic diagram of the intelligent question-answering structure based on multi-round network retrieval and self-correction provided by the present invention.
[0049] Appendix Figure 6 This is a schematic diagram of the multi-round network retrieval self-correction device provided by the present invention.
[0050] Appendix Figure 7 This is a schematic diagram of the intelligent question-answering device based on multi-round network retrieval self-correction provided by the present invention. Detailed Implementation
[0051] The present invention is not limited to the following embodiments, and the specific implementation can be determined according to the technical solution of the present invention and the actual situation.
[0052] Those skilled in the art will understand that, unless specifically stated otherwise, in the embodiments of the present invention, a "module" or "unit" refers to a computer program or part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0053] In addition, in the embodiments of the present invention, "multiple" refers to two or more, and "first" and "second" are used to distinguish descriptions and should not be construed as implying relative importance.
[0054] This invention provides a multi-round network retrieval self-correction method, an intelligent question-answering method based on multi-round network retrieval self-correction, and related apparatus. The intelligent question-answering apparatus based on multi-round network retrieval self-correction can be integrated into a computer device, which can be a server, a terminal, or other similar device; it can also be executed jointly by a terminal and a server. The above examples should not be construed as limiting the invention.
[0055] The aforementioned terminals may include mobile phones, wearable smart devices, tablet computers, laptops, personal computers (PCs), and in-vehicle computers, etc., and this invention does not limit them. This invention also does not limit the number of terminal devices.
[0056] The aforementioned server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This invention does not limit these features.
[0057] For example, a computer device performs similarity matching between a given question and a preliminary retrieved answer to obtain a similarity level. The preliminary answer includes one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and dissimilarity. In response to high similarity, a multi-round network retrieval self-correction process is not triggered. In response to similarity levels other than high similarity, different rounds of network retrieval self-correction processes are triggered according to a preset triggering mechanism. The given question and knowledge blocks are input into the question-answering model to obtain the final answer to the given question, where the question-answering model is based on a large language model.
[0058] Example 1: As shown in the attached document Figure 1 As shown, this embodiment of the invention discloses a multi-round network retrieval self-correction method, including:
[0059] Step S110: Perform similarity matching between the given question and the retrieved preliminary answer to obtain a similarity level. The preliminary answer includes one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and no similarity.
[0060] Step S120: In response to high similarity, the multi-round network retrieval self-correction process is not triggered;
[0061] Step S130: In response to similarity levels other than high similarity, trigger different rounds of network retrieval self-correction process according to a preset triggering mechanism, wherein the preset triggering mechanism includes:
[0062] If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results.
[0063] If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time.
[0064] If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
[0065] In step S110 above, a similarity match is performed between the given question and the retrieved preliminary answer to obtain a similarity level. This level determines whether the retrieved preliminary answer is sufficient to meet the requirements of answering the given question. The various similarity levels can be labeled in different ways. For example, "Similar" can be marked in green; "Low Similarity" can be marked in yellow; "Middle" can be marked in blue; and "Unlikeness" can be marked in red.
[0066] If the similarity level in step S120 above is high similarity, it means that the preliminary answer retrieved is sufficient to meet the solution requirements of the given question. In this case, no additional network search is required, and the preliminary answer retrieved from the knowledge base is directly output.
[0067] If the similarity level in step S130 above is low, it indicates that the preliminary answer retrieved has a low relevance to the given question. To improve the accuracy and relevance of the knowledge, a network search (i.e., the first network search) is needed. The network search results are then used to correct the preliminary answer for the first time, and the knowledge block obtained from the first correction is output. Specifically, the network search involves calling the network search interface W to retrieve the knowledge block C = {d} related to the given question X. c1 ,d c2 ...d ck}
[0068] If the similarity level in step S130 above is medium, it means that although the preliminary answer retrieved is somewhat related to the given question, it lacks comprehensiveness and directness. In order to obtain more comprehensive and direct knowledge, based on the first network search, the given question is decomposed into sub-questions, and network searches are performed for each sub-question. Then, all network search results are used to perform a second correction on the first correction result, and the knowledge block obtained from the second correction is output.
[0069] If the similarity level in step S130 above is dissimilar, it means that the preliminary answer retrieved is completely unrelated to the question. In this case, in order to enhance the relevance and accuracy of knowledge, HYDE (Hypothetical Document Embedding Algorithm) is used to perform network retrieval to retrieve knowledge blocks related to the second network retrieval. The network retrieval results are then used to perform a third correction on the second correction results, and the knowledge blocks obtained from the third correction are output.
[0070] The formulaic expression of steps S120 to S130 above is as follows:
[0071] ①Label = "Similar"
[0072] Output(KB(Q))
[0073] ②Label="LowSimilarity"
[0074] WebSearch1(Q)→C(KB(Q)) i →Output(C(KB(Q)))
[0075] ③Label = "Middle"
[0076]
[0077] {WebSearch2(SubQ i )KB(Q)→Output(C(KB(Q)) i )
[0078] ④Label = “Unlikeness”
[0079]
[0080] HYDESearch(Q,{WebSearch2(SubQ i )})→C(KB(Q)) i →Output(C(KB(Q)))
[0081] Where Q represents the user's given question; KB(Q) represents the knowledge retrieved from the knowledge base related to the given question Q; Sim(KB(Q),Q) represents the similarity between the knowledge base retrieval result KB(Q) and the given question Q; Label represents the label determined based on the similarity; WebSearch1(Q) represents the result of the first web search; SubQi represents the i-th sub-question after decomposing the given question Q; C(KB(Q)) i : Indicates that the i-th knowledge block is corrected; WebSearch2(SubQ i ): Represents the result of the second web search performed on the sub-problem SubQi; HYDESearch(Q,WebSearch2): Represents the result of a more in-depth web search performed using the HYDE algorithm.
[0082] This invention discloses a multi-round network retrieval self-correction method. By matching the similarity between a given question and the preliminary retrieved answer, different rounds of network retrieval self-correction processes are triggered. This greatly improves the accuracy and efficiency of obtaining relevant knowledge blocks through network retrieval. Furthermore, the multi-round network retrieval self-correction improves the relevance between the retrieved knowledge blocks and the question, which is conducive to improving the question-answering accuracy of subsequent large language models.
[0083] Example 2: As shown in the attached document Figure 2 As shown, this embodiment of the invention discloses a multi-round network retrieval self-correction method, including:
[0084] Step S210, obtaining the given question and the retrieved preliminary answer, including:
[0085] (1) Obtain the questions raised by users and use them as given questions;
[0086] (2) Retrieve one or more knowledge blocks most relevant to the question from the sentence-level RAG knowledge base to form a preliminary answer.
[0087] In question-answering, the construction of the RAG knowledge base plays a crucial role in knowledge retrieval for large language models, with the knowledge segmentation strategy within the knowledge base being particularly important. Knowledge segmentation in the knowledge base mainly falls into three categories: character-length-based segmentation, sentence- or paragraph-based segmentation, etc. However, the chunking obtained from segmenting the original document using these methods often suffers from semantic incompleteness or knowledge redundancy, which may mislead the large language model's responses. (Appendix) Figure 3 This relates to the impact of knowledge segmentation on the model's response. Semantically incomplete knowledge is likely to mislead the responses of large language models. Therefore, a more precise knowledge segmentation strategy is needed to segment out more semantically complete knowledge to guide large language models in outputting accurate answers.
[0088] Therefore, this embodiment addresses the problem of semantic damage caused by improper document segmentation strategies in traditional RAG knowledge base construction by setting up a sentence-level RAG knowledge base optimization strategy. Specifically, firstly, the documents to be processed are segmented into sentences to reduce their fineness. Less fine-grained knowledge has more complete semantics after being merged through mergibility analysis. Secondly, the sliding window method is used to judge the mergibility of two adjacent sentences in combination with analysis indicators, including semantic coherence, semantic completeness, and semantic clarity. "Semantic coherence" and "semantic completeness" can ensure that the two sentences have more complete and coherent semantics after merging, while "semantic clarity" can ensure that the two sentences have clearer and more explicit semantics after merging. When the three conditions are met, adjacent sentences are merged. The merged chunking semantic information is richer and does not contain redundant information, which can provide more accurate knowledge prompts for large language models.
[0089] Step S220: Perform similarity matching on the given question and the retrieved preliminary answer to obtain a similarity level, where the similarity level includes high similarity, low similarity, medium similarity, and no similarity.
[0090] In this embodiment, similarity matching between a given question and the retrieved preliminary answer is performed using a similarity tag generation model. The construction process of the similarity tag generation model includes:
[0091] (1) Use the ChatGPT 3.5 model to generate a fine-tuning dataset, which includes several fine-tuning data. Each fine-tuning data includes a question, the knowledge blocks that the question may retrieve, and the similarity between the question and the knowledge blocks.
[0092] (2) Supervised fine-tuning training of the pre-trained model is performed using the fine-tuning dataset to obtain the similarity label generation model, wherein the pre-trained model is a small model obtained by training using the training dataset.
[0093] The similarity label generation model construction process in this embodiment aims to effectively transfer the advantages of the high-performance language model ChatGPT 3.5 to a locally deployable, small-scale model. Specifically, it utilizes the ChatGPT 3.5 model to generate a high-quality dataset containing (question-knowledge-similarity). This dataset is designed to capture complex language patterns and semantic relationships, providing a solid foundation for subsequent model training. Crucially, the ChatGPT 3.5 model, with its powerful language understanding and generation capabilities, can accurately identify and associate questions with relevant knowledge, while assigning them appropriate similarity labels, thereby ensuring the accuracy and richness of the dataset.
[0094] Furthermore, the Prompt template can be used to generate a fine-tuned dataset using the GPT3.5 model. The Prompt template not only covers various language scenarios and semantic relationships but also fully considers the diversity and complexity of the problems to ensure that the generated dataset comprehensively reflects real-world language usage. This approach provides rich and diverse data support for training the local model, further enhancing its generalization ability and adaptability. The dataset construction using the Prompt template with the GPT3.5 model is shown in Table 1.
[0095] Table 1 shows the dataset construction using the Prompt template with the GPT 3.5 model.
[0096]
[0097] In this embodiment, a locally deployable small model is trained using Supervised Fine-Tuning (SFT) on a fine-tuning dataset. SFT, as an effective transfer learning method, allows the extraction of valuable knowledge and information from large pre-trained models and its "injection" into smaller, resource-constrained models, thereby significantly improving their performance without sacrificing deployment convenience. The goal of SFT training in this embodiment is to enable the local model to inherit the ChatGPT 3.5 model's superior ability in similarity label generation, that is, to accurately and efficiently determine the degree of similarity between a problem and its related knowledge.
[0098] Step S230: In response to high similarity, the multi-round network retrieval self-correction process is not triggered.
[0099] Step S240: In response to similarity levels other than high similarity, trigger different rounds of network retrieval self-correction process according to a preset triggering mechanism, wherein the preset triggering mechanism includes:
[0100] If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results.
[0101] If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time.
[0102] If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
[0103] In this embodiment of the invention, a sentence-level optimization strategy is designed to optimize the knowledge base structure, ensuring the integrity of the knowledge base structure. Furthermore, a (question-knowledge-similarity) Prompt template is designed, and a fine-tuning dataset is constructed using the ChatGPT3.5 model. The pre-trained model is then supervised and fine-tuned using this dataset to obtain a similarity label generation model. Based on this model, the similarity matching result between a given question and the retrieved preliminary answer is determined, thereby triggering different rounds of network retrieval self-correction processes. This improves the accuracy and efficiency of acquiring relevant knowledge blocks through network retrieval, ultimately enhancing the question-answering accuracy of the subsequent large language model.
[0104] Example 3: As shown in the attached document Figure 4 , 5 As shown, this embodiment of the invention discloses an intelligent question-answering method based on multi-round network retrieval self-correction, including:
[0105] Step S310: Obtain the given problem and obtain the knowledge block of the given problem through a multi-round network retrieval self-correction method;
[0106] It should be noted that the knowledge block obtained by the multi-round network retrieval self-correction method disclosed in Examples 1 and 2 can be a preliminary answer retrieved from the knowledge base or a knowledge block obtained by different rounds of network retrieval self-correction.
[0107] Step S320: Input the given question and knowledge block into the question answering model to obtain the final answer to the given question, wherein the question answering model is based on the large language model.
[0108] The question-answering models include:
[0109] A word embedding model vectorizes a given question and knowledge block as input; here, the word embedding model can be referred to as a word embedding model.
[0110] The re-ranking model re-orders knowledge blocks and uses the Sort function to score the similarity between knowledge blocks and a given question, selecting the Top_K knowledge blocks with the highest similarity. Here, the BGE Reranker-large model can be used to re-rank knowledge blocks retrieved from the knowledge base and from the web.
[0111] The large language model analyzes a given question, integrates knowledge blocks, and generates the final answer.
[0112] The formalized expression of the large language model is as follows:
[0113] Answer = LLM({Q}∪{D} i |i∈[1,n]})
[0114] Where Answer represents the final generated answer, LLM represents the Large Language Model, {Q} represents the user's initial question being a text set, and D... i |i∈[1,n]} represents n precise knowledge blocks obtained through network retrieval.
[0115] The knowledge and the given question are input into a large language model. Leveraging the powerful text understanding and generation capabilities of the large language model, the content is automatically integrated to generate the final answer. The formulaic expression of the large language model is as follows:
[0116] Answer = LLM({Q}U{D) i |i∈[1,n]})
[0117] Where Answer represents the final generated answer, LLM represents the Large Language Model, {Q} represents the user's given question as a text set, and D... i|i∈[1,n]} represents n precise knowledge blocks obtained through knowledge base or network retrieval.
[0118] Example 4: Generalization experiments were conducted using single-hop, multi-hop, and triplet datasets. Data sets from three vertical fields—finance, law, and medicine—were selected to verify the cross-domain generalization ability of this invention, as detailed below:
[0119] (I) Creating a dataset
[0120] The SQuAD2.0 single-round public question-answering dataset contains natural questions, including queries and related documents containing answers. This dataset is used to verify the capabilities of multi-round network retrieval self-correction methods for single-round intelligent question answering.
[0121] The HotpotQA multi-turn public question-answering dataset contains questions and their multi-turn dialogues (including answers) related to social hot topics. This dataset is used to verify the capabilities of the multi-turn network retrieval self-correction method for intelligent question answering in multi-turn dialogues.
[0122] PopQA triplet public question-answering dataset, with data formatted as triplet (question-background knowledge-answer) natural world questions. This dataset is used to verify the ability of multi-round network retrieval self-correction methods to understand and transform background knowledge.
[0123] The vertical domain datasets selected are Chinese-law-and-regulations, FinanceIQ, and mmlu_professional_medicine. Domain generalization experiments were conducted on the model in these three vertical domains to verify the performance of the Ir-Web-Cot intelligent question answering framework constructed by integrating the multi-round network retrieval self-correction method and sentence-level RAG knowledge base partitioning strategy in this embodiment.
[0124] (II) Evaluation Indicators
[0125] The final output results of this research method were comprehensively evaluated using machine learning metrics (Accuracy, Recall, and F1-Score) and the ChatGPT3.5 large language model. The specific evaluation content of the metrics is shown in Table 2.
[0126] Table 2 Evaluation Indicators and Specific Evaluation Content
[0127]
[0128] (III) Model Selection
[0129] The web retrieval model adopts the commercial version of ZhiPuAI's GLM-4-Air web retrieval model, which integrates the web retrieval interface into GLM-4-Air.
[0130] For the word embedding model, we chose the open-source word embedding model bge-large-zh-v1.5 from BAAI (Beijing Academy of Artificial Intelligence), which is trained using the RetroMAE method.
[0131] The re-ranking model uses the BGE Reranker-large model to re-rank the knowledge retrieved from the knowledge base and the network, and uses the Sort function to score the similarity between knowledge and questions. Finally, the Top_K (K=6) contexts with the highest similarity are selected.
[0132] For the large language model, the mainstream open-source model Tongyi Qianwen Qwen2.5 was chosen. To facilitate local deployment and ensure performance, versions Qwen2.5:7b and Qwen2.5:14b were selected. This model uses 18T tokens for pre-training, supports context lengths of 4096 and text generation of up to 8000 characters, accurately understands user questions, and generates Chinese results. This helps the large language model fully understand the context and output accurate answers.
[0133] (IV) Results Analysis
[0134] The baseline, standard Cot, Ir_Cot, and Ir_Web_Cot (the multi-round network retrieval self-correction method of this invention) were compared. Experiments were conducted on three public datasets, SQuAD2.0, HotpotQA, and PopQA, using two models from the Qianwen series, Qwen2.5:7b and Qwen2.5:14b, and three basic models, ChatGPT3.5-Turbo. The specific experimental results are shown in Table 3.
[0135]
[0136] Secondly, to evaluate the overall performance of Ir-Web-RAG (an intelligent question-answering model framework built by integrating Ir_Web_Cot with the RAG system), this embodiment selected C-RAG, Adaptive-RAG, and Self-RAG as comparison methods and conducted comparative experiments on three public benchmark datasets (SQuAD2.0, HotpotQA, and PopQA). In addition, two models from the Qianwen series (Qwen2.5-7B and Qwen2.5-14B) and ChatGPT-3.5-Turbo were selected as baseline models for comparison. Detailed experimental results are shown in Table 4.
[0137] Table 4 Experimental Results
[0138]
[0139] Finally, relevant experiments were conducted in vertical fields such as finance, law, and healthcare to verify the generalization ability of the Ir-Web-RAG model. The experimental results are shown in Table 5.
[0140] Table 5 Experimental Results
[0141]
[0142] As shown in Table 3, the experimental results demonstrate that this invention outperforms other models in terms of accuracy, recall, and F1-score, based on the three large language models Qwen2.5:7b, Qwen2.5:14b, and ChatGPT3.5-Turbo. The Ir-cot method adds multiple rounds of knowledge base retrieval to the traditional Cot method, repeatedly correcting the retrieved knowledge blocks to improve accuracy and relevance. Ir-Web-Cot integrates the knowledge base into the network, fine-tuning the question-knowledge similarity label generation model to improve the efficiency of network knowledge retrieval. It leverages the high timeliness and accuracy of network retrieval to perform multiple rounds of error correction on the knowledge blocks retrieved from the knowledge base, thereby improving the model's answer accuracy.
[0143] Table 4 shows that constructing Ir-Web-RAG can provide more relevant and accurate knowledge to large language models, resulting in more accurate results. Optimizing the sentence-level RAG knowledge base improves the semantic completeness of the knowledge, and more complete knowledge can better guide the output of large language models. Combining web retrieval with the RAG system overcomes the closed nature of traditional RAG knowledge bases, improving the accuracy and relevance of professional knowledge.
[0144] Finally, generalization experiments were conducted in the fields of finance, law, and medicine to demonstrate the powerful generalization ability of the Ir-Web-RAG invention, as shown in Table 5. However, multiple rounds of network retrieval inevitably lead to longer model response time. By slightly sacrificing response time, the accuracy of the model's response can be greatly increased, thereby greatly reducing the possibility of hallucinatory responses and ensuring the credibility of the large language model's answers to questions.
[0145] (V) Evaluation of Large Language Models
[0146] ChatGPT3.5 was provided with a series of questions, preset target evaluation metrics, and a pair of candidate answers. ChatGPT3.5's task was to compare and analyze the two answers based on these metrics to determine which answer performed better under the given metrics, and to explain the basis for its judgment. During the evaluation process, the following principles were followed: if one answer was significantly better than the other, that answer was considered the winner; if both performed similarly and the difference was within an acceptable range, it was considered a tie. To reduce the impact of randomness that might be introduced by large language models (LLMs) when processing tasks, this embodiment adopted a strategy of averaging multiple runs. Specifically, for each pair of answers, we repeated the evaluation five times, and finally determined the final result based on the average score of the five evaluations.
[0147] To visually demonstrate the effectiveness of this evaluation method, Table 6 provides a specific evaluation example, showcasing the practical application of LLM in the evaluation process and its resulting evaluation results. This approach allows for a more scientific and objective evaluation of the generated Ir_Web_RAG.
[0148] Table 6. Overall Evaluation Results
[0149]
[0150]
[0151] Example 5: As shown in the attached document Figure 6 As shown, this embodiment of the invention discloses a multi-round network retrieval self-correction device, comprising:
[0152] The similarity level determination unit performs similarity matching between a given question and the retrieved preliminary answer to obtain a similarity level. The preliminary answer includes one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and no similarity.
[0153] Network retrieval routing unit, including:
[0154] In response to high similarity, it does not trigger a multi-round network retrieval self-correction process;
[0155] In response to similarity levels other than high similarity, different rounds of network retrieval self-correction processes are triggered according to preset triggering mechanisms, including:
[0156] If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results.
[0157] If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time.
[0158] If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
[0159] The multi-round network retrieval self-correction device in this embodiment integrates multi-round network retrieval self-correction and sentence-level RAG knowledge base partitioning strategies to construct the Ir-Web-Cot intelligent question answering framework. It performs similarity matching between a given question and the retrieved preliminary answer, and selects corresponding rounds of network retrieval strategies for different similarity levels to perform information retrieval and error correction, greatly improving the accuracy and efficiency of acquiring relevant knowledge through network retrieval. The specific steps of each unit in this embodiment are the same as in the above embodiments and will not be repeated.
[0160] Example 6: As attached Figure 7 As shown, this embodiment of the invention discloses an intelligent question-answering device based on multi-round network retrieval self-correction, comprising:
[0161] The basic data acquisition unit acquires a given question and obtains knowledge blocks for the given question through a multi-round network retrieval self-correction method;
[0162] The question-answering analysis unit takes a given question and knowledge block as input to the question-answering model and obtains the final answer to the given question. The question-answering model is based on a large language model.
[0163] The specific steps of each unit in this embodiment are the same as those in the above embodiments, and will not be repeated here.
[0164] The intelligent question-answering device based on multi-round network retrieval self-correction disclosed in this embodiment integrates the RAG framework (the basic structure of question-answering analysis units) with Ir-Web-Cot to construct the Ir-Web-RAG intelligent question-answering framework, which can improve the timeliness and relevance of knowledge block acquisition, thereby prompting the large language model to provide accurate answers to user questions.
[0165] Example 7: This embodiment of the invention discloses a storage medium storing a computer program that can be read by a computer. The computer program is configured to execute an intelligent question answering method based on multi-round network retrieval self-correction when running.
[0166] The aforementioned storage media may include, but are not limited to, USB flash drives, read-only memory, portable hard drives, magnetic disks, optical disks, and other media capable of storing computer programs.
[0167] Example 8: This embodiment of the invention discloses an electronic device, including a processor and a memory. The memory stores a computer program, which is loaded and executed by the processor to implement an intelligent question-answering method based on multi-round network retrieval self-correction.
[0168] The processor described above can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. It can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The memory can include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, portable hard drives, magnetic disks, or optical disks.
[0169] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0170] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0171] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.
[0172] The above content is only a specific embodiment of the present invention, which has strong adaptability and implementation effect. However, the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be covered within the protection scope of the present invention. Therefore, equivalent changes made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A multi-round network retrieval self-correction method, characterized in that, include: A similarity match is performed between a given question and the retrieved preliminary answers to obtain a similarity level. The preliminary answers include one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and no similarity. In response to high similarity, it does not trigger a multi-round network retrieval self-correction process; In response to similarity levels other than high similarity, different rounds of network retrieval self-correction processes are triggered according to preset triggering mechanisms, including: If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results. If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time. If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
2. The multi-round network retrieval self-correction method according to claim 1, characterized in that, The similarity matching between the given question and the retrieved answer is performed using a similarity tag generation model, wherein the construction process of the similarity tag generation model includes: A fine-tuning dataset was generated using the ChatGPT 3.5 model. The fine-tuning dataset includes several modeling datasets, each of which includes a question, the knowledge blocks that the question may retrieve, and the similarity between the question and the knowledge blocks. Supervised fine-tuning training of the pre-trained model is performed using the fine-tuning dataset to obtain a similarity label generation model, where the pre-trained model is a small model obtained by training using the training dataset.
3. The multi-round network retrieval self-correction method according to claim 1 or 2, characterized in that, It also includes obtaining a given question and the preliminary answers retrieved, including: Obtain the questions submitted by users and use them as given questions; Retrieve one or more knowledge blocks most relevant to the question from the sentence-level RAG knowledge base to form a preliminary answer.
4. The multi-round network retrieval self-correction method according to claim 3, characterized in that, The optimization strategies for the sentence-level RAG knowledge base include: The documents to be processed are divided into sentence-level segments; The sliding window method is used to determine the merging of two adjacent sentences by combining analytical indicators, including semantic coherence, semantic integrity, and semantic clarity.
5. An intelligent question-answering method based on multi-round network retrieval self-correction, characterized in that, include: Obtain a given problem and obtain a knowledge block of the given problem using the method described in any one of claims 1 to 4; Input a given question and knowledge block into the question answering model to get the final answer to the given question. The question answering model is based on a large language model.
6. The intelligent question-answering method based on multi-round network retrieval self-correction according to claim 5, characterized in that, The question-answering model includes: Word embedding models vectorize the given input question and knowledge block; The reordering model reorders knowledge blocks and uses the Sort function to score the similarity between the knowledge blocks and the given question, selecting the Top_K knowledge blocks with the highest similarity. The large language model analyzes a given question, integrates knowledge blocks, and generates the final answer.
7. A multi-round network retrieval self-correction device applying the method described in any one of claims 1 to 4, characterized in that, include: The similarity level determination unit performs similarity matching between a given question and the retrieved preliminary answer to obtain a similarity level. The preliminary answer includes one or more knowledge blocks in the knowledge base that are most relevant to the given question. The similarity levels include high similarity, low similarity, medium similarity, and no similarity. Network retrieval routing unit, including: In response to high similarity, it does not trigger a multi-round network retrieval self-correction process; In response to similarity levels other than high similarity, different rounds of network retrieval self-correction processes are triggered according to preset triggering mechanisms, including: If the similarity level is low, then a first online search is performed, and the preliminary answer is corrected for the first time using the search results. If the similarity level is medium, a second network search is performed. The problem is decomposed into sub-problems, and a network search is performed on each sub-problem. The network search results are then used to correct the first error correction result a second time. If the similarity level is not similar, a third network search is performed using the HYDE algorithm. The network search results are then used to correct the second error correction result a third time.
8. An intelligent question-answering device based on multi-round network retrieval self-correction, applying the method as described in any one of claims 5 to 6, characterized in that, include: A basic data acquisition unit acquires a given problem and obtains a knowledge block of the given problem using the method described in any one of claims 1 to 4. The question-answering analysis unit takes a given question and knowledge block as input to the question-answering model and obtains the final answer to the given question. The question-answering model is based on a large language model.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the steps of the method as described in any one of claims 5 to 6.
10. A storage medium, characterized in that, The storage medium stores a computer program that can be read by a computer, the computer program being configured to execute the steps of the method as described in any one of claims 5 to 6 when it is run.