Intelligent question answering method and system based on multi-module collaborative optimization

By adopting multi-module collaborative optimization methods in the intelligent question-answer system, including knowledge scope judgment, dynamic retrieval, multi-level problem rewriting, knowledge screening and self-reflection optimization modules, the existing system's shortcomings in resource consumption, complex problem handling and field adaptability are solved, and efficient and accurate intelligent question-and-answer effects are achieved.

CN119557409BActive Publication Date: 2025-05-09SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +4
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510121723.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-09
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

现有智能问答系统在资源消耗、复杂问题处理及领域适应性方面存在不足,且查询重写模块生成的单一查询可能限制信息多样性,输入问题与底层查询意图之间的偏差导致无法准确解读用户需求。

Method used

An intelligent question-answer method based on multi-module collaborative optimization is adopted, including knowledge scope judgment module, dynamic search module, multi-level question rewriting module, knowledge screening module and self-reflection optimization module. Through the collaborative work of these modules, a complete intelligent question-answer link is formed, semantic expansion and decomposition are realized, and the relevance and context adaptation capabilities of the search results are improved.

Benefits of technology

It significantly improves the adaptability and answer quality of the intelligent question-and-answer system in complex scenarios, reduces computing resource consumption, improves the training efficiency of the model and domain adaptability, and ensures the semantic integrity and logical rigor of the output content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557409B_ABST
    Figure CN119557409B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of knowledge question answering, and to an intelligent question answering method and system based on multi-module collaborative optimization. The method comprises: inputting a question to be answered into a knowledge question answering model, and the knowledge question answering model outputs a knowledge question answering result; a knowledge scope judgment module in the model judges whether the problem can be solved by relying on its own knowledge, and if not, enters a dynamic retrieval module; the dynamic retrieval module performs similarity retrieval on the content of a memory knowledge base according to the question to be answered, and if the retrieval result does not meet the requirements, enters a multi-level question rewriting module; the multi-level question rewriting module rewrites the question to be answered, and inputs the rewritten question into a knowledge screening module; the knowledge screening module outputs a screened document according to the rewritten question, and a self-reflection optimization module generates a preliminary answer according to the document and the question, and judges whether the preliminary answer is reasonable, and if it is unreasonable, performs self-reflection optimization, thereby providing a new solution for the development of intelligent question answering technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge question answering, and in particular to an intelligent question answering method and system based on multi-module collaborative optimization. Background Art

[0002] Large language models (LLMs) perform well in knowledge reasoning and multi-task processing, but the parameterized knowledge stored internally may have limitations and is difficult to quickly update or integrate the latest information. To address this problem, the retrieval-augmented generation (RAG) method was proposed to dynamically extract non-parametric knowledge through the retrieval of external knowledge and documents. This method greatly improves the accuracy and adaptability of LLMs in contextual problems.

[0003] The basic RAG system consists of a knowledge retrieval module and a reading module, forming a retrieval-reading process. However, this basic pipeline has problems such as low retrieval efficiency and insufficient reliability of generated answers. To address these problems, the RAG framework has gradually incorporated more complex modular designs. For example, the query rewriting module, as a connector between the input question and the retrieval module, optimizes the retrieval effect by rewriting the query, thus evolving a rewrite-retrieval-reading architecture. At the same time, models such as RETA-LLM and RARR add post-reading and fact-checking modules to improve the accuracy and reliability of answers. In addition, the addition of modules such as query routers and resource rankers further enhances the applicability of RAG in complex scenarios. The integration of these modules ultimately forms the paradigm of modular RAG, transforming the traditional RAG process into a highly flexible and dynamic system.

[0004] Although modular RAG has made significant progress, there are still some problems in its practical application. For example, query rewriting modules usually generate a single query, which may limit the diversity of retrievable information. In addition, the deviation between the input question and the underlying query intent, especially the intent misalignment caused by ambiguous expressions, often leads to an inaccurate interpretation of user needs. Although query rewriting can improve the retrieval of relevant information, it cannot guarantee the accuracy of the information. Overly broad retrieval may also introduce noise and affect the quality of the answer. Summary of the invention

[0005] In order to solve the deficiencies of the prior art, the present invention provides an intelligent question-answering method and system based on multi-module collaborative optimization;

[0006] On the one hand, an intelligent question-answering method based on multi-module collaborative optimization is provided, including:

[0007] Input the question to be answered into the knowledge question answering model, and the knowledge question answering model outputs the knowledge question answering result;

[0008] Among them, the knowledge scope judgment module in the knowledge question answering model judges whether the problem can be solved by relying on its own knowledge. If it can, it enters the self-reflection optimization module, and if it cannot, it enters the dynamic retrieval module; the acquisition process of the knowledge scope judgment module includes: selecting the Llama 3B model; constructing the first training set; using the first training set to train the Llama 3B model, when the weighted dynamic correction loss function value no longer decreases, stop training, and obtain the trained Llama 3B model; the trained Llama 3B model is used as the knowledge scope judgment module;

[0009] The weighted dynamic correction loss function is specifically expressed as follows:

[0010] ;

[0011] Where N is the total number of samples, For sample The true label of is 0 or 1; Samples predicted by the model The probability of belonging to category 1; is the sample weight; is the dynamic correction coefficient, which is used to control the intensity of error correction; is the error sensitivity index, which adjusts the response of the correction term to the error amplitude. represents the weighted dynamic correction loss function;

[0012] The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered;

[0013] The multi-level question rewriting module rewrites the questions to be answered and inputs the rewritten questions into the knowledge screening module;

[0014] The knowledge screening module outputs the screened documents based on the rewritten questions. If the number of documents exceeds zero, the documents and questions are input into the self-reflection optimization module;

[0015] The self-reflection optimization module generates a preliminary answer based on the document and the question, and determines whether the preliminary answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed; the self-reflection optimization module includes: inputting the original question, rewritten question and screened document into the large language model to obtain the preliminary generated answer; performing self-reflection check on the preliminary generated answer through the large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

[0016] On the other hand, an intelligent question-answering system based on multi-module collaborative optimization is provided, including:

[0017] The question-answering module is configured to: input the question to be answered into the knowledge question-answering model, and the knowledge question-answering model outputs the knowledge question-answering result;

[0018] Among them, the knowledge scope judgment module in the knowledge question answering model judges whether the problem can be solved by relying on its own knowledge. If it can, it enters the self-reflection optimization module, and if it cannot, it enters the dynamic retrieval module; the acquisition process of the knowledge scope judgment module includes: selecting the Llama 3B model; constructing the first training set; using the first training set to train the Llama 3B model, when the weighted dynamic correction loss function value no longer decreases, stop training, and obtain the trained Llama 3B model; the trained Llama 3B model is used as the knowledge scope judgment module;

[0019] The weighted dynamic correction loss function is specifically expressed as follows:

[0020] ;

[0021] Where N is the total number of samples, For sample The true label of is 0 or 1; Samples predicted by the model The probability of belonging to category 1; is the sample weight; is the dynamic correction coefficient, which is used to control the intensity of error correction; is the error sensitivity index, which adjusts the response of the correction term to the error amplitude. represents the weighted dynamic correction loss function;

[0022] The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered;

[0023] The multi-level question rewriting module rewrites the questions to be answered and inputs the rewritten questions into the knowledge screening module;

[0024] The knowledge screening module outputs the screened documents based on the rewritten questions. If the number of documents exceeds zero, the documents and questions are input into the self-reflection optimization module;

[0025] The self-reflection optimization module generates a preliminary answer based on the document and the question, and determines whether the preliminary answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed; the self-reflection optimization module includes: inputting the original question, rewritten question and screened document into the large language model to obtain the preliminary generated answer; performing self-reflection check on the preliminary generated answer through the large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

[0026] The above technical solution has the following advantages or beneficial effects:

[0027] Aiming at the performance improvement and multi-domain adaptation requirements of the intelligent question-answering system, the present invention proposes an extensible modular question-answering system and construction method based on a small parameter model drive. Through the improved fine-tuning method and module collaborative architecture, many shortcomings of traditional methods in resource consumption, complex problem processing and domain adaptability are effectively solved. The system adopts a multi-layer modular design, which not only optimizes the full-link process from knowledge retrieval to answer generation, but also significantly improves the efficiency and adaptability of the system in multiple scenarios and multiple tasks, providing a new solution for the development of intelligent question-answering technology.

[0028] This paper introduces a fine-tuning method based on hierarchical freezing and selective parameter activation in a small parameter model. By dynamically activating the key parameter layer related to the task, it avoids updating all parameters, significantly reduces computing resource consumption, and improves the training efficiency and domain adaptability of the model. In addition, combined with the designed weighted dynamic correction loss function (WDCL), it shows excellent performance in dealing with class imbalance problems and complex sample optimization, and significantly improves the response accuracy of the model in multiple scenarios.

[0029] The present invention forms a complete intelligent question-answering link from knowledge judgment to question generation to answer optimization by designing multiple modules to work together. The designed dynamic retrieval and multi-level question rewriting module work together to achieve semantic expansion and decomposition by converting the questions input by the user into a set of sub-questions with both coarse and fine granularity, improve the relevance and context adaptation ability of the retrieval results, and ensure the reliability of the results in the retrieval stage through a dynamic confidence scoring mechanism; the self-reflection optimization module dynamically improves the generated answers through logical consistency checks and multiple rounds of verification to ensure the semantic integrity and logical rigor of the output content, especially in complex field problems. The knowledge scope judgment module gives priority to using internal knowledge to reduce dependence on external resources, combined with the dynamic update and historical reuse capabilities of the memory knowledge base module, greatly improves the system response speed and long-term self-learning ability, and significantly enhances the efficiency and adaptability of the intelligent question-answering system.

[0030] Based on the modular design concept, the present invention divides the intelligent question-answering process into multiple modules with independent functions and high collaboration, including knowledge scope judgment, dynamic retrieval, multi-level question rewriting, knowledge screening, self-reflection optimization and memory knowledge base modules. Compared with the fixed retrieval-reading process in the traditional RAG (retrieval enhancement generation) method, the present invention realizes a more flexible and dynamic architecture through modular design. Each module realizes data flow and task collaboration through a standardized interface, which not only improves the overall operation efficiency of the system, but also significantly enhances the adaptability of the system in different scenarios and user needs. The traditional RAG method is easily limited to a single retrieval path and fixed generation logic when dealing with complex problems. The present invention effectively solves the semantic deviation and inaccurate answer problems in complex domain problems through the synergy of multi-level question rewriting and self-reflection optimization modules. At the same time, the division of labor and collaboration between modules support multi-task parallel processing, significantly improve system performance and user experience, and provide powerful intelligent support capabilities for multi-domain applications.

[0031] In summary, the present invention significantly improves the adaptability and answer quality of the intelligent question-answering system in complex scenarios through the combination of modular design, multi-level question rewriting and dynamic optimization mechanism. Compared with the traditional RAG method, the present invention shows a more flexible, dynamic and efficient architectural advantage, not only making breakthroughs in retrieval accuracy and generation logic, but also significantly leading in domain scalability and system efficiency, providing accurate, efficient and reliable technical support for multi-domain intelligent question-answering applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0033] Figure 1 This is a schematic diagram of the internal connection relationship of the knowledge question and answer model of Example 1. DETAILED DESCRIPTION

[0034] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] Embodiment 1

[0037] This embodiment provides an intelligent question-answering method based on multi-module collaborative optimization;

[0038] Intelligent question-answering method based on multi-module collaborative optimization, including:

[0039] Input the question to be answered into the knowledge question answering model, and the knowledge question answering model outputs the knowledge question answering result;

[0040] Among them, the knowledge scope judgment module in the knowledge question answering model judges whether the problem can be solved by relying on its own knowledge. If it can, it enters the self-reflection optimization module, and if it cannot, it enters the dynamic retrieval module; the acquisition process of the knowledge scope judgment module includes: selecting the Llama 3B model; constructing the first training set; using the first training set to train the Llama 3B model, when the weighted dynamic correction loss function value no longer decreases, stop training, and obtain the trained Llama 3B model; the trained Llama 3B model is used as the knowledge scope judgment module;

[0041] The weighted dynamic correction loss function is specifically expressed as follows:

[0042] ;

[0043] Where N is the total number of samples, For sample The true label of is 0 or 1; Samples predicted by the model The probability of belonging to category 1; is the sample weight; is the dynamic correction coefficient, which is used to control the intensity of error correction; is the error sensitivity index, which adjusts the response of the correction term to the error amplitude. represents the weighted dynamic correction loss function;

[0044] The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered;

[0045] The multi-level question rewriting module rewrites the questions to be answered and inputs the rewritten questions into the knowledge screening module;

[0046] The knowledge screening module outputs the screened documents based on the rewritten questions. If the number of documents exceeds zero, the documents and questions are input into the self-reflection optimization module;

[0047] The self-reflection optimization module generates a preliminary answer based on the document and the question, and determines whether the preliminary answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed; the self-reflection optimization module includes: inputting the original question, rewritten question and screened document into the large language model to obtain the preliminary generated answer; performing self-reflection check on the preliminary generated answer through the large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

[0048] Furthermore, if Figure 1 As shown, the knowledge question and answer model includes: a knowledge scope judgment module, a dynamic retrieval module, a multi-level question rewriting module, a knowledge screening module and a self-reflection optimization module which are connected in sequence; the output end of the knowledge screening module is also connected to the input end of the multi-level question rewriting module; the input end of the self-reflection optimization module is also connected to the output end of the knowledge scope judgment module; the input end of the dynamic retrieval module is also connected to the memory knowledge base.

[0049] Furthermore, the knowledge scope judgment module is used to vectorize the question to be answered to obtain the question vector to be answered, compare the similarity between the question vector to be answered and the knowledge inside the knowledge scope judgment module, judge whether the current question belongs to the knowledge scope learned by the knowledge scope judgment module, and obtain the question category to which the question to be answered belongs; the question category includes: problems that can be solved based on existing knowledge or problems that cannot be solved based on existing knowledge.

[0050] Furthermore, constructing the first training set includes:

[0051] (1-1): Randomly select M question and correct answer pairs from the public dataset;

[0052] (1-2): Use GPT-4 to give M reasoning answers to M questions;

[0053] (1-3): Compare the inference answer corresponding to each question with the correct answer to see if they are consistent. If they are consistent, the current question is considered to be a problem that can be solved based on existing knowledge, and the current question is stored in the self-knowledge set; if they are inconsistent, the current question is considered to be a problem that cannot be solved based on existing knowledge, and the current question is stored in the unknown knowledge set;

[0054] (1-4): Use the self-knowledge set and the unknown knowledge set as the first training set.

[0055] Furthermore, the Llama 3B model is trained using the first training set, and only some key parameter layers are adjusted during the training process, and a selective parameter update strategy is also introduced.

[0056] The key parameter layers include: output layer, attention modules from the 28th to the 32nd layers, feedforward networks FFN from the 28th to the 32nd layers, and embedding layers:

[0057] The output layer is located at the top of the model and is responsible for mapping the output of the feedforward network to the output distribution of the target task. As the core layer that directly controls the task output, optimizing the output layer can quickly improve the accuracy and logic of the generated results.

[0058] The attention modules from the 28th to the 32nd layers capture the high-order semantic associations between input and context through a multi-head self-attention mechanism. They mainly optimize the projection matrices of Query, Key, and Value (responsible for semantic matching and relationship modeling) and the output transformation matrix (optimizing the feature mapping of multi-head attention), significantly improving the accuracy of semantic modeling and the quality of answer generation.

[0059] The feed-forward network FFN from the 28th to the 32nd layer performs nonlinear transformation and feature expansion on the output of the attention module. Optimizing the weights of the nonlinear activation part of the FFN from the 28th to the 32nd layer can enhance the model's adaptability to complex semantic features.

[0060] The embedding layer maps discrete inputs into high-dimensional semantic vectors. When targeting domain tasks, the embedding layer can be optimized slightly to enhance the understanding of domain vocabulary, but it is usually kept frozen in general tasks.

[0061] The selective parameter update strategy specifically includes: combining the improved weighted dynamic correction loss function (WDCL) to dynamically amplify the impact of high-error samples, ensuring that the optimization process focuses on the critical parameter layers (such as the attention modules and output layers in the last few layers). By freezing the first 27 layers of Transformer blocks, only the attention modules from the 28th to the 32nd layers and the core parts of FFN (such as the projection matrices of Query, Key and Value, nonlinear activation weights, etc.) are updated, reducing unnecessary computational overhead and improving the performance of the model in complex tasks.

[0062] It should be understood that when a user inputs a question, the system will first pass the question to the trained knowledge scope judgment module. The knowledge scope judgment module uses a fine-tuned small model based on the previously constructed data set to quickly identify whether the question falls within the scope of the model's own knowledge coverage. This process relies on the model's internal knowledge representation capabilities, and through efficient vector matching and semantic understanding technology, it achieves accurate classification of problems. If the module determines that the question can be answered using the model's own knowledge, there is no need to call on additional resources, and the system will directly enter the self-reflection optimization stage. If the module determines that the question raised by the user exceeds the coverage of the model's internal knowledge (that is, it belongs to ), the system will automatically trigger the dynamic retrieval module. This module prioritizes using the model's existing knowledge to quickly solve problems, and only calls external resources to supplement when the model knowledge is insufficient. This strategy not only saves computing resources, but also significantly improves the efficiency of answering, especially for scenarios with high requirements for response speed. At the same time, the two paths of self-reflection optimization and dynamic retrieval complement each other, ensuring the system's answer quality and knowledge expansion capabilities, and providing users with an intelligent question-and-answer service experience that is both efficient and reliable.

[0063] It should be understood that when making the knowledge scope judgment module fine-tuning dataset, we first randomly extract 1,000 high-quality question-answer pairs from each public QA dataset (Natural Question (NQ), TriviaQA, StrategyQA, HotpotQA, 2WikiMQA) to form a question set 。 During the data screening process, manual review and authoritative verification are introduced to ensure that the selected questions are accurate and the answers are credible.

[0064] It should be understood that in the data generation stage, a large language model L is used to target each question. Reasoning to generate answers , forming a generated data set During the generation process, the model output is optimized by designing prompt templates. For example, (1) the model is guided to answer questions based on its own knowledge. The template can be written as "Please answer the following questions based on your knowledge: {question}, answer". (2) A few-shot learning strategy is introduced to improve the model's ability to understand the context of the question by providing a small number of example questions and answers. For complex questions, context enhancement technology is used to add relevant background prompts to the questions to improve the model's reasoning ability.

[0065] It should be understood that after generating the answer, the model's generated answer With the correct answer If the two are consistent, it is considered that the model can answer questions based on its own knowledge and such questions are classified as self-knowledge set. ; If the two are inconsistent, they are classified into the unknown knowledge set Through this type of classification, a fine-tuned dataset is formed In this way, the model's knowledge coverage and capability boundaries are divided as much as possible.

[0066] It should be understood that after the dataset is constructed, the small model is fine-tuned by an improved fine-tuning method. In response to the problem of computing resource consumption during model fine-tuning, a fine-tuning method based on layered freezing and parameter activation is designed. By freezing the backbone network of the small model Llama 3B, only some key parameter layers are adjusted. The designed method introduces a selective parameter update strategy, which adaptively selects parameters in the network that are more relevant to the target task for adjustment, thereby avoiding unnecessary gradient updates, further saving computing resources, and improving the training efficiency and generalization performance of the model. In order to enhance the robustness of the fine-tuning process to different types of problems, a new loss function is designed, called the Weighted Dynamic Correction Loss (WDCL). The designed loss function not only considers the problem of category imbalance, but also dynamically adjusts the error weights of positive and negative examples to ensure the accuracy of the model's answers to questions in different scenarios. The designed loss function adds a dynamic error correction mechanism on the basis of the traditional weighted cross entropy loss. By correcting the term Adaptively adjust the attention to difficult samples. The loss function comprehensively considers the sample weights and prediction error amplitude, achieving dual optimization for class imbalance and difficult-to-handle samples. Compared with existing methods (such as weighted cross entropy and focal loss), WDCL can significantly improve the model's response accuracy and robustness to complex samples while ensuring the stability of easy-to-classify samples. This method can not only reduce the computational overhead during fine-tuning, but also is suitable for application scenarios with limited data scale and complex and diverse tasks, demonstrating strong innovation and potential application value.

[0067] It should be understood that a user feedback mechanism is introduced to support users to evaluate the answer results, and the problems that can be solved will be recorded by the system. According to the data set The module is updated regularly based on the update status. The self-evolution capability of the submodule can be realized to continuously improve the solution of complex problems.

[0068] It should be understood that a user feedback mechanism is introduced to support users to evaluate the answer results, and unresolved issues will be recorded by the system. According to the data set The module is updated regularly to keep up with the latest updates. The module's self-evolution capability is achieved, which can continuously improve the solutions to complex problems.

[0069] Furthermore, the dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered, and outputs the result if the retrieval result meets the requirements, and enters the multi-level question rewriting module if the retrieval result does not meet the requirements. The dynamic retrieval module includes:

[0070] The question to be answered is converted into a semantic vector, and the similarity between the semantic vector of the question to be answered and the semantic vector of the historical question in the memory knowledge base is calculated based on the approximate nearest neighbor algorithm of L2 distance, and several candidate retrieval results with the highest similarity are selected; L2 represents the Euclidean distance;

[0071] A confidence score is calculated for each candidate retrieval result and compared with a dynamic confidence threshold: if the confidence score is higher than the threshold, the candidate retrieval result is retained and used as the answer to the question to be answered; otherwise, the current candidate retrieval result is passed to the next step.

[0072] Furthermore, the L2 distance-based approximate nearest neighbor algorithm calculates the similarity between the semantic vector of the question to be answered and the semantic vector of the historical question in the memory knowledge base, including:

[0073] ;

[0074] in, is the vector representation of the input problem, is the vector representation of historical questions in the knowledge base, is the dimension of the vector. Returns several candidate results with the smallest distance and their corresponding metadata. Represents the similarity between the semantic vector of the question to be answered and the semantic vector of the historical questions in the memory knowledge base.

[0075] Furthermore, calculating a confidence score for each candidate search result includes:

[0076] ;

[0077] in, is the mean of the current search result set, is the standard deviation of the current search result set, is the similarity between the current query vector and the result vector, represents the confidence score;

[0078] Furthermore, the dynamic confidence threshold includes:

[0079] The threshold is dynamically set according to the confidence score to decide whether to use the current result directly or to conduct further search and optimization. The dynamic confidence threshold formula is:

[0080] ;

[0081] Represents the dynamic confidence threshold, which is used to determine whether the current search result is credible;

[0082] It represents the mean of the similarity scores in the current search result set, indicating the similarity level of the overall results;

[0083] Represents the standard deviation of the similarity scores in the current search result set, measuring the degree of dispersion of the similarity distribution;

[0084] It represents adjustment parameters, which are dynamically adjusted based on user behavior or historical query data, reflecting the flexibility of confidence requirements.

[0085] It should be understood that the dynamic retrieval module, after receiving the user input question, converts the question into a high-dimensional semantic vector through the Chinese embedding model. The semantic vector of the user question is queried through the Milvus search interface, and the similarity with the historical questions in the knowledge base is quickly calculated using the approximate nearest neighbor (ANN) algorithm based on L2 distance.

[0086] It should be understood that after converting the user question into a vector representation, the similarity between the query vector and the historical question vectors in the knowledge base is calculated. In order to enhance the accuracy and robustness of the system results, an adaptive confidence scoring mechanism is introduced to dynamically decide whether to use the results directly or to search further based on the confidence of the query results.

[0087] The confidence score of each result is dynamically calculated according to the formula to measure the reliability of the current result.

[0088] It should be understood that if the confidence score is above the dynamic confidence threshold , the system directly returns the current result. If it is lower than the threshold, the system triggers the subsequent steps, such as passing it to the multi-level question rewriting module to perform the next step.

[0089] It should be understood that after receiving the question input by the user, the dynamic retrieval module first converts the question into a high-dimensional semantic vector through the Chinese embedding model, and uses the Milvus vector database combined with the approximate nearest neighbor (ANN) algorithm based on L2 distance to calculate the similarity between the input vector and the historical question vector, and quickly screens out several candidate results with the smallest distance. On this basis, the system introduces a designed adaptive confidence scoring mechanism, which dynamically evaluates the confidence score of each result according to the formula by calculating the mean and standard deviation of the retrieval result set to measure the reliability of the current result. The system screens the retrieval results according to the dynamic confidence threshold: if the confidence score is higher than the dynamic threshold, the current result is directly returned as the answer; if it is lower than the threshold, the system passes the question to the multi-level question rewriting module for further optimization. This design logic ensures that the module can not only efficiently process inputs similar to historical questions, but also intelligently adapt to new questions, improving the response efficiency and robustness of the system.

[0090] Furthermore, the multi-level question rewriting module rewrites the question to be answered, wherein the multi-level question rewriting module includes:

[0091] Constructing a second data set, the second data set comprising: an original question, a set of coarse-grained questions corresponding to the original question, and a set of fine-grained questions corresponding to the original question;

[0092] The small model Gemma 2B is adopted and the second training set is used to train the small model Gemma 2B to obtain the trained small model Gemma 2B-Rewriter. During the training process, the input value of the small model Gemma 2B is the original problem, and the output value of the small model Gemma 2B is the coarse-grained problem set corresponding to the original problem and the fine-grained problem set corresponding to the original problem.

[0093] Furthermore, the constructing of the second data set specifically includes:

[0094] Get the original question, use GPT-4 to rewrite the original question in a coarse-grained manner, and obtain a set of coarse-grained questions;

[0095] For each coarse-grained question, GPT-4 is used to perform fine-grained rewriting to obtain a set of fine-grained questions.

[0096] The coarse-grained rewriting includes: breaking down a problem into several sub-problems, replacing or fuzzily expressing polysemous words, or introducing domain knowledge.

[0097] The fine-grained rewriting includes: adding context-related modifiers, qualifiers or parameterized options to the question.

[0098] It should be understood that the multi-level question rewriting module is driven by a pre-trained small model, receives questions passed from the dynamic retrieval module, rewrites them into a set of sub-questions with both coarse-grained and fine-grained levels, and sends these sub-questions to the next module for retrieval or further processing. The specific logic is: the module receives questions passed by the retrieval trigger module or other modules as input. The small model first generates a number of coarse-grained rewritten questions, and then further refines each coarse-grained question to generate multiple refined sub-questions. All generated coarse-grained and fine-grained sub-questions are passed to the downstream knowledge screening module to complete the information query. For "unmatched questions" returned from the knowledge screening module, the module will try to rewrite them. If the number of rewrites exceeds three times and a suitable answer cannot be obtained, the output "Cannot solve the problem" will be directly output, the problem will be recorded and the user will be prompted to optimize the question.

[0099] The multi-level question rewriting module first relies on GPT-4 to generate a structured rewriting dataset. The dataset consists of 3,000 pieces of data collected from the above five datasets. Through the designed question rewriting prompt, the model is guided to perform multi-level semantic expansion of user questions and generate rewriting results from coarse-grained to fine-grained. The question rewriting prompt example is as follows:

[0100] [Instruction] Perform two levels of optimization on the original problem:

[0101] 1. Coarse-grained rewriting: Convert the original question into a concise and clear statement, keeping the semantics intact but making the language more fluent and clear.

[0102] 2. Fine-grained expansion: Based on the rewritten questions, further explore the diverse semantic levels and specific intentions to generate more detailed versions of the questions.

[0103] [Examples]

[0104] Here are some examples to illustrate the format and requirements:

[0105] [Example 1]

[0106] Original question: how do quantum computers solve optimization problems. Translation: How do quantum computers solve optimization problems.

[0107] Coarse-grained rewrite:

[0108] How do quantum computers approach solving optimization problems. Translation: How do quantum computers approach solving optimization problems.

[0109] Fine-grained extensions:

[0110] 1. What techniques do quantum computers use to solve optimization challenges. Translation: What techniques do quantum computers use to solve optimization challenges.

[0111] 2. How do quantum algorithms outperform classical methods in optimization tasks. Translation: How quantum algorithms outperform classical methods in optimization tasks.

[0112] 3. What are the key principles behind quantum computing's effectiveness in optimization. Translation: What are the key principles behind the efficiency of quantum computing in optimization problems.

[0113] 4. Which types of optimization problems are best suited for quantum computing. Translation: Which types of optimization problems are best suited for quantum computing.

[0114] 5. What are real-world examples of quantum computers solving optimization problems. Translation: What are real-world examples of quantum computers solving optimization problems.

[0115] Coarse-grained rewriting focuses on the global rewriting of the problem, aiming to expand the semantic scope and reconstruct the original problem from multiple perspectives. Specifically, it includes: 1. Simplifying complex problems into several sub-problems, such as decomposing "how to optimize the performance of distributed databases on blockchain" into "how to improve the throughput of blockchain" and "how to reduce the query latency of distributed databases". 2. Replacing polysemous words or fuzzy expressions to generate queries with different semantic versions, for example, "distributed" can be replaced with "sharding" or "multi-node architecture". 3. By introducing domain knowledge, it captures the possible implicit requirements of the problem and generates multiple semantic variants.

[0116] On the basis of coarse-grained rewriting, further refinement and expansion are carried out for each rewriting problem, adding more context-related modifiers, qualifiers or parameterized options. For example, "How to optimize blockchain performance" can be refined into: "How to optimize blockchain TPS (transactions per second)", "How to balance node load in sharded blockchain", "How to reduce communication delay of blockchain network".

[0117] Each refined question focuses on a different sub-field or technical solution to improve the retrieval accuracy of the question. This multi-level expansion strategy from coarse to fine enables the system to both broaden the scope of question coverage and conduct in-depth mining for specific information needs.

[0118] The questions generated by the above strategy are used to rewrite the dataset. The dataset is in the form of:

[0119]

[0120] in, is a single piece of data after decomposition in the data set. For the original question, is the set of coarse-grained rewriting problems, A collection of questions that extend to fine-grained levels.

[0121] This dataset is used to fine-tune small models (such as Gemma 2B). The fine-tuned model has stronger multi-level question rewriting capabilities and can accurately handle complex queries in the field.

[0122] In order to achieve high-performance, low-latency online services, the module uses the FastAPI framework to build a lightweight service interface and integrates a fine-tuned rewrite model.

[0123] Furthermore, the knowledge screening module includes:

[0124] Constructing a third data set, wherein the third data set includes: a question set, a document set, an explanation, and a classification result;

[0125] Explanation is a logical description of the classification result, which is used to explain why a specific classification result is obtained. It is not equivalent to the answer to the question set, but provides a verifiable semantic association description for the classification decision. This explanation can provide the reason for the model classification result and help verify whether the classification decision is reasonable. It shows users or system developers how the model obtains the classification result, thereby enhancing the transparency and trust of the system. The explanation information can be used by subsequent modules (such as self-reflection optimization module or generation module) to improve the logic and coherence of the generated answer.

[0126] For example: "Question": "What is blockchain".

[0127] "Document": [ "1. Blockchain is a distributed ledger technology that ensures the security and immutability of data through encryption algorithms.", "2. Blockchain ensures the security and immutability of data through encryption algorithms. Each block contains a set of transaction records and is linked to the previous block through a hash value.", "3. The core feature of blockchain is decentralization. Data is stored on multiple nodes in the network rather than concentrated on a central server."].

[0128] "Explanation": [ "1. It comprehensively covers the core features of blockchain (distributed ledger, encryption algorithm, decentralization, hash link) and application areas, which is highly consistent with the knowledge context, so it is classified as "relevant.", "2. It accurately describes the core features and technical principles of blockchain, which is consistent with the knowledge context, so it is classified as "relevant.", "3. It is partially correct, but does not cover the extended information of blockchain's encryption algorithm, hash link and application areas, so it is classified as "neutral."].

[0129] "Classification results": ["1. Relevant","2. Relevant","3. Neutral"].

[0130] Input the third data set into the Llama 3B model, train the Llama 3B model, and obtain a trained Llama 3B-Filter model;

[0131] During the training process, the question set is used as the input value of the Llama 3B model, and the document set, explanation and classification results are used as the output value of the Llama 3B model.

[0132] Furthermore, constructing the third data set includes:

[0133] Input the question set into the search engine to obtain documents containing the correct answer and documents not containing the correct answer;

[0134] Input the question set and the obtained document set into GPT-4 to obtain the classification results and the corresponding explanations of the classification results.

[0135] It should be understood that the knowledge screening module receives the decomposed question set s at the input stage. Each sub-question will serve as the core of the retrieval, triggering the retrieval of relevant documents k. The retrieved documents then enter the processing stage, where they are evaluated in conjunction with the GPT model to generate an explanation e and classification result r for each document.

[0136] Documents are classified into the categories of "relevant", "neutral", or "irrelevant" depending on their relevance to the question.

[0137] If the document is marked as "relevant", it proceeds directly to the next stage.

[0138] If the document is "neutral", choose to keep it or continue to decompose the problem for optimization based on specific needs.

[0139] For "irrelevant" documents, delete this search document.

[0140] In the output phase, the system will return the retained documents and interpretation results.

[0141] If the number of relevant documents in the remaining documents is greater than 0, it will go to the next module; if there are no relevant documents, the sub-question will be rewritten. The system's update mechanism is dynamic and can adjust and expand the data set in a timely manner as the tasks change and the continuous feedback of user interactions. Through regular model retraining and data set updates, the model will continue to learn and adapt to new task requirements, maintaining its efficiency and accuracy.

[0142] Use the decomposed problem set , and the documents retrieved by the search engine as context , as a prompt for GPT-4 to generate a short explanation And the classification results For the decomposed problem set ,based on Retrieved knowledge , the module evaluates the retrieved knowledge Whether it contains information that is helpful in answering the question. Therefore, the corresponding classification results will be generated. and interpretation of the results The classification results include three types: relevant, neutral and irrelevant.

[0143] For the knowledge that is relevant in the classification results, the relevant knowledge is retained; for the irrelevant knowledge, the retreat strategy is adopted to combine the questions Further decomposition is performed; neutral knowledge is retained or further decomposed according to the characteristics of the processing task. Finally, a data set is formed .

[0144] The constructed dataset is used to fine-tune the Llama 3B model. The training combines the generation task with the classification task, and optimizes the model through the loss function so that it can learn two core capabilities: one is to accurately generate explanations based on documents. , and secondly, effectively assess the relevance of the document to the question This training method not only improves the reasoning ability of the model, but also significantly reduces the use of computing resources, so that the model can not only classify document relevance, but also provide reasonable explanations to meet the needs of complex task scenarios. The fine-tuned Llama 3B-Filter model is deployed as the core of this function, and the FastAPI framework is used to build a lightweight service interface.

[0145] Furthermore, the self-reflection optimization module includes:

[0146] Input the original question, rewritten question, and screened documents into the large language model to obtain a preliminary generated answer;

[0147] The initially generated answers are self-reflectively checked through a large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

[0148] The self-reflective optimization module first receives three input data: the original question, the rewritten question, and the filtered related documents. The rewritten question is an analysis or rewriting of the original question, the purpose of which is to make the question easier to understand or clearer. The filtered documents are the relevant contextual information, which is used to assist in generating more accurate answers. The model generates a preliminary answer based on the rewritten question and the filtered document content. In order to ensure that the generated answer meets the requirements of the question, the module designs a self-reflective prompt word prompt, that is, by checking whether the generated answer is logically clear, consistent with the document content, and whether it fully answers the question. If the model finds that the generated answer has logical inconsistencies, omissions or errors, it will guide the model to adjust the generation method through a feedback mechanism and regenerate it. This process will be repeated within a limited maximum number of cycles, usually 3 to 5 times. If the model still cannot generate a reasonable answer within the specified number of times, the module will output "Cannot solve this problem" and end the process, and the relevant questions will be recorded in the database. On the contrary, if the generated answer meets the requirements, the module will immediately return the final answer.

[0149] The self-reflection optimization module accepts the original question input by the user, the rewritten question set, and the filtered related document content. The original question may be relatively general or complex, and the rewritten question set is an optimization of the original question to make it more specific and clear. The filtered related documents contain information such as background knowledge, context or professional terms related to the question, which helps the model generate accurate answers.

[0150] The self-reflective optimization module generates preliminary answers based on the rewritten question set and filtered documents. At this stage, the model generates a response that is as complete as possible by combining the semantics of the question and the content of the document. The answer at this point may be an inferred answer based on the information in the document and has not yet been quality-checked.

[0151] The self-reflection optimization module uses a set of prompts to allow the model to self-check the generated answers. These prompts include: answer consistency: whether the generated answer is consistent with the information in the document; answer completeness: whether the answer covers all the key information of the question; logic check: whether the logic of the answer is coherent and there are no obvious errors or contradictions; semantic accuracy: whether the generated answer accurately reflects the meaning of the question.

[0152] If the self-reflection optimization module finds problems with the answer during the self-reflection process (such as missing information, unclear semantics or logical errors, etc.), the model will modify it according to the feedback and regenerate the answer. This process can be completed by adjusting the generation strategy or optimizing the prompt word. The adjusted generation will enter the next round of inspection to ensure that each round of generation is more accurate than the previous round.

[0153] The self-reflection optimization module sets a maximum number of cycles (usually 3 to 5 times), repeatedly generates and checks the quality of the answer within this number limit. Each iteration improves the generated content based on the feedback from the previous iteration until the answer meets the quality requirements or the maximum number of iterations is reached.

[0154] If the answer finally generated by the model meets expectations within the maximum number of iterations, the self-reflection optimization module directly outputs the answer. If the model still cannot give a satisfactory answer after multiple iterations, it outputs "cannot solve this problem" and provides reasons or suggests further operations.

[0155] The memory knowledge base stores the questions that the system has successfully answered and their corresponding answers in the form of vectors, and provides an efficient retrieval and dynamic update mechanism. By storing historical questions and answers as high-dimensional semantic vectors, the module can quickly match the relevance of the user's current questions with known questions, thereby avoiding repeated retrieval and improving response efficiency. At the same time, the memory knowledge base block combines access frequency and data update time to dynamically clean up redundant data and maintain the efficiency and scalability of the knowledge base. The memory knowledge base not only reduces the system's dependence on external resources, but also provides basic support for the system's self-learning and long-term optimization, significantly enhancing the adaptability and intelligence of the question-answering system in multiple scenarios.

[0156] The memory knowledge base collects historical answer data of the system and stores questions, answers, and documents after retrieval and screening as basic data. It removes redundant, repeated or noisy data and standardizes the expression of questions (such as sentence conversion and spelling check) to unify the semantic format.

[0157] The memory knowledge base uses a pre-trained model to convert each question into a high-dimensional semantic vector, retaining its semantic features. The embedded vector of the question and the corresponding metadata (answer text, timestamp, relevance score, etc.) are stored together to provide support for subsequent retrieval and dynamic management.

[0158] The memory knowledge base uses the Milvus vector database to build an efficient index structure to support approximate nearest neighbor (ANN) search of large-scale vectors.

[0159] The natural language questions input by the user are passed to the knowledge scope judgment module through the API.

[0160] The knowledge scope judgment module makes judgments based on its own knowledge. If the problem can be solved by relying on the model’s own knowledge, it enters the self-reflection optimization module; otherwise, it enters the dynamic retrieval module.

[0161] The dynamic retrieval module performs rapid similarity retrieval on the knowledge base and decides whether to return the results or pass them to the question rewriting module based on the retrieval results and the set retrieval threshold.

[0162] The multi-level question rewriting module reconstructs the questions and then passes the question set to the knowledge screening module.

[0163] The knowledge screening module retrieves the question set to obtain relevant documents, and screens according to the questions and related documents. If the number of supporting documents is greater than 0, it enters the self-reflection optimization module. If there is no supporting document, it will return to the multi-level question rewriting module to further decompose the question set.

[0164] The self-reflection optimization module first generates a preliminary answer based on the screened documents and question set, and determines whether the generated answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed again.

[0165] The memory knowledge base module stores and learns all valid records in the data stream for subsequent knowledge updating and model optimization.

[0166] Embodiment 2

[0167] This embodiment provides an intelligent question-answering system based on multi-module collaborative optimization, including:

[0168] The question-answering module is configured to: input the question to be answered into the knowledge question-answering model, and the knowledge question-answering model outputs the knowledge question-answering result;

[0169] Among them, the knowledge scope judgment module in the knowledge question answering model judges whether the problem can be solved by relying on its own knowledge. If it can, it enters the self-reflection optimization module, and if it cannot, it enters the dynamic retrieval module; the acquisition process of the knowledge scope judgment module includes: selecting the Llama 3B model; constructing the first training set; using the first training set to train the Llama 3B model, when the weighted dynamic correction loss function value no longer decreases, stop training, and obtain the trained Llama 3B model; the trained Llama 3B model is used as the knowledge scope judgment module;

[0170] The weighted dynamic correction loss function is specifically expressed as follows:

[0171] ;

[0172] Where N is the total number of samples, For sample The true label of is 0 or 1; Samples predicted by the model The probability of belonging to category 1; is the sample weight; is the dynamic correction coefficient, which is used to control the intensity of error correction; is the error sensitivity index, which adjusts the response of the correction term to the error amplitude. represents the weighted dynamic correction loss function;

[0173] The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered;

[0174] The multi-level question rewriting module rewrites the questions to be answered and inputs the rewritten questions into the knowledge screening module;

[0175] The knowledge screening module outputs the screened documents based on the rewritten questions. If the number of documents exceeds zero, the documents and questions are input into the self-reflection optimization module;

[0176] The self-reflection optimization module generates a preliminary answer based on the document and the question, and determines whether the preliminary answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed; the self-reflection optimization module includes: inputting the original question, rewritten question and screened document into the large language model to obtain the preliminary generated answer; performing self-reflection check on the preliminary generated answer through the large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

[0177] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent question-answering method based on multi-module collaborative optimization, characterized by: include: Input the question to be answered into the knowledge question answering model, and the knowledge question answering model outputs the knowledge question answering result; Among them, the knowledge scope judgment module in the knowledge question answering model judges whether the problem can be solved by relying on its own knowledge. If it can, it enters the self-reflection optimization module, and if it cannot, it enters the dynamic retrieval module; the acquisition process of the knowledge scope judgment module includes: selecting the Llama 3B model; constructing the first training set; using the first training set to train the Llama 3B model, when the weighted dynamic correction loss function value no longer decreases, stop training, and obtain the trained Llama 3B model; the trained Llama 3B model is used as the knowledge scope judgment module; The first training set is used to train the Llama 3B model, and only some key parameter layers are adjusted during the training process, and a selective parameter update strategy is also introduced; The key parameter layers specifically include: output layer, attention modules from the 28th to the 32nd layers, feedforward networks FFN from the 28th to the 32nd layers, and embedding layers; The weighted dynamic correction loss function is specifically expressed as follows: ; Where N is the total number of samples, For sample The true label of is 0 or 1; Samples predicted by the model The probability of belonging to category 1; is the sample weight; is the dynamic correction coefficient, which is used to control the intensity of error correction; is the error sensitivity index, which adjusts the response of the correction term to the error amplitude. represents the weighted dynamic correction loss function; The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered; Dynamic confidence threshold, including: dynamically setting the threshold according to the confidence score to decide whether to directly use the current result or to conduct further search and optimization; the dynamic confidence threshold formula is: ; Represents the dynamic confidence threshold, which is used to determine whether the current search result is credible; It represents the mean of the similarity scores in the current search result set, indicating the similarity level of the overall results; Represents the standard deviation of the similarity scores in the current search result set, measuring the degree of dispersion of the similarity distribution; It represents adjustment parameters, which are dynamically adjusted based on user behavior or historical query data, reflecting the flexibility of confidence requirements; The multi-level question rewriting module rewrites the questions to be answered and inputs the rewritten questions into the knowledge screening module; The multi-level question rewriting module rewrites the questions to be answered, wherein the multi-level question rewriting module includes: Constructing a second data set, the second data set comprising: an original question, a set of coarse-grained questions corresponding to the original question, and a set of fine-grained questions corresponding to the original question; The small model Gemma 2B is used, and the second training set is used to train the small model Gemma 2B to obtain the trained small model Gemma 2B-Rewriter; during the training process, the input value of the small model Gemma 2B is the original problem, and the output value of the small model Gemma2B is the coarse-grained problem set corresponding to the original problem and the fine-grained problem set corresponding to the original problem; The knowledge screening module outputs the screened documents based on the rewritten questions. If the number of documents exceeds zero, the documents and questions are input into the self-reflection optimization module; The knowledge screening module constructs a third data set, including a question set, a document set, an explanation, and a classification result; The knowledge screening module receives the decomposed question set s at the input stage, and each sub-question will serve as the core of the retrieval, triggering the retrieval of relevant documents k; the retrieved documents then enter the processing stage, and are evaluated in combination with the GPT model to generate the explanation e and classification result r of each document; Documents will be classified as "relevant", "neutral", or "irrelevant" depending on their relevance to the question; If the document is marked as "relevant", it goes directly to the next stage; If the document is "neutral", choose to keep it or continue to decompose the problem for optimization based on specific needs; For "irrelevant" documents, delete this search document; In the output phase, the system will return the retained documents and interpretation results; The self-reflection optimization module generates a preliminary answer based on the document and the question, and determines whether the preliminary answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed; the self-reflection optimization module includes: inputting the original question, rewritten question and screened document into the large language model to obtain the preliminary generated answer; performing self-reflection check on the preliminary generated answer through the large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

2. The intelligent question-answering method based on multi-module collaborative optimization according to claim 1, characterized in that: The knowledge scope judgment module is used to vectorize the question to be answered to obtain the question vector to be answered, compare the question vector to be answered with the knowledge inside the knowledge scope judgment module for similarity, judge whether the current question belongs to the knowledge scope learned by the knowledge scope judgment module, and obtain the question category to which the question to be answered belongs; The problem categories include: problems that can be solved based on existing knowledge or problems that cannot be solved based on existing knowledge.

3. The intelligent question-answering method based on multi-module collaborative optimization according to claim 1, characterized in that: The step of constructing a first training set comprises: Randomly select M question-answer pairs from the public dataset; Use GPT-4 to give M reasoning answers to M questions; Compare whether the inference answer corresponding to each question is consistent with the correct answer. If they are consistent, it is considered that the current problem is a problem that can be solved based on the existing knowledge, and the current problem is stored in the self-knowledge set; if they are inconsistent, it is considered that the current problem is a problem that cannot be solved based on the existing knowledge, and the current problem is stored in the unknown knowledge set; the self-knowledge set and the unknown knowledge set are used as the first training set.

4. The intelligent question-answering method based on multi-module collaborative optimization according to claim 1, characterized in that: The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered. The dynamic retrieval module includes: The question to be answered is converted into a semantic vector, and the similarity between the semantic vector of the question to be answered and the semantic vector of the historical question in the memory knowledge base is calculated based on the approximate nearest neighbor algorithm of L2 distance, and several candidate retrieval results with the highest similarity are selected; L2 represents the Euclidean distance; A confidence score is calculated for each candidate retrieval result and compared with a dynamic confidence threshold: if the confidence score is higher than the threshold, the candidate retrieval result is retained and used as the answer to the question to be answered; otherwise, the current candidate retrieval result is passed to the next step.

5. The intelligent question-answering method based on multi-module collaborative optimization according to claim 4, characterized in that: The L2 distance-based approximate nearest neighbor algorithm calculates the similarity between the semantic vector of the question to be answered and the semantic vector of the historical question in the memory knowledge base, including: ; in, is the vector representation of the input problem, is the vector representation of historical questions in the knowledge base, is the dimension of the vector; returns several candidate results with the smallest distance and their corresponding metadata; Represents the similarity between the semantic vector of the question to be answered and the semantic vector of the historical questions in the memory knowledge base; The step of calculating a confidence score for each candidate search result includes: ; in, is the mean of the current search result set, is the standard deviation of the current search result set, is the similarity between the current query vector and the result vector; Represents the confidence score.

6. The intelligent question-answering method based on multi-module collaborative optimization according to claim 1, characterized in that: The constructing of the second data set specifically includes: Get the original question, use GPT-4 to rewrite it in a coarse-grained manner to obtain a set of coarse-grained questions; for each coarse-grained question, use GPT-4 to rewrite it in a fine-grained manner to obtain a set of fine-grained questions; The coarse-grained rewriting includes: breaking down a problem into several sub-problems, replacing or fuzzifying polysemous words, or introducing domain knowledge; The fine-grained rewriting includes: adding context-related modifiers, qualifiers or parameterized options to the question.

7. The intelligent question-answering method based on multi-module collaborative optimization according to claim 1, characterized in that: The knowledge screening module comprises: Constructing a third data set, wherein the third data set includes: a question set, a document set, an explanation, and a classification result; Input the third data set into the Llama 3B model, train the Llama 3B model, and obtain a trained Llama 3B-Filter model; During the training process, the question set is used as the input value of the Llama 3B model, and the document set, explanation and classification results are used as the output value of the Llama 3B model.

8. The intelligent question-answering method based on multi-module collaborative optimization according to claim 7, characterized in that: The constructing of the third data set includes: inputting the question set into a search engine to obtain documents containing correct answers and documents not containing correct answers; Input the question set and the obtained document set into GPT-4 to obtain the classification results and the corresponding explanations of the classification results.

9. The intelligent question-answering system based on multi-module collaborative optimization is characterized by: include: The question-answering module is configured to: input the question to be answered into the knowledge question-answering model, and the knowledge question-answering model outputs the knowledge question-answering result; Among them, the knowledge scope judgment module in the knowledge question answering model judges whether the problem can be solved by relying on its own knowledge. If it can, it enters the self-reflection optimization module, and if it cannot, it enters the dynamic retrieval module; the acquisition process of the knowledge scope judgment module includes: selecting the Llama 3B model; constructing the first training set; using the first training set to train the Llama 3B model, when the weighted dynamic correction loss function value no longer decreases, stop training, and obtain the trained Llama 3B model; the trained Llama 3B model is used as the knowledge scope judgment module; The first training set is used to train the Llama 3B model, and only some key parameter layers are adjusted during the training process, and a selective parameter update strategy is also introduced; The key parameter layers specifically include: output layer, attention modules from the 28th to the 32nd layers, feedforward networks FFN from the 28th to the 32nd layers, and embedding layers; The weighted dynamic correction loss function is specifically expressed as follows: ; Where N is the total number of samples, For sample The true label of is 0 or 1; Samples predicted by the model The probability of belonging to category 1; is the sample weight; is the dynamic correction coefficient, which is used to control the intensity of error correction; is the error sensitivity index, which adjusts the response of the correction term to the error amplitude. represents the weighted dynamic correction loss function; The dynamic retrieval module performs similarity retrieval on the content of the memory knowledge base according to the question to be answered. If the retrieval result meets the requirements, the result is output. If the retrieval result does not meet the requirements, the multi-level question rewriting module is entered; Dynamic confidence threshold, including: dynamically setting the threshold according to the confidence score to decide whether to directly use the current result or to conduct further search and optimization; the dynamic confidence threshold formula is: ; Represents the dynamic confidence threshold, which is used to determine whether the current search result is credible; It represents the mean of the similarity scores in the current search result set, indicating the similarity level of the overall results; Represents the standard deviation of the similarity scores in the current search result set, measuring the degree of dispersion of the similarity distribution; It represents adjustment parameters, which are dynamically adjusted based on user behavior or historical query data, reflecting the flexibility of confidence requirements; The multi-level question rewriting module rewrites the questions to be answered and inputs the rewritten questions into the knowledge screening module; The multi-level question rewriting module rewrites the questions to be answered, wherein the multi-level question rewriting module includes: Constructing a second data set, the second data set comprising: an original question, a set of coarse-grained questions corresponding to the original question, and a set of fine-grained questions corresponding to the original question; The small model Gemma 2B is used, and the second training set is used to train the small model Gemma 2B to obtain the trained small model Gemma 2B-Rewriter; during the training process, the input value of the small model Gemma 2B is the original problem, and the output value of the small model Gemma2B is the coarse-grained problem set corresponding to the original problem and the fine-grained problem set corresponding to the original problem; The knowledge screening module outputs the screened documents based on the rewritten questions. If the number of documents exceeds zero, the documents and questions are input into the self-reflection optimization module; The knowledge screening module constructs a third data set, including a question set, a document set, an explanation, and a classification result; The knowledge screening module receives the decomposed question set s at the input stage, and each sub-question will serve as the core of the retrieval, triggering the retrieval of relevant documents k; the retrieved documents then enter the processing stage, and are evaluated in combination with the GPT model to generate the explanation e and classification result r of each document; Documents will be classified as "relevant", "neutral", or "irrelevant" depending on their relevance to the question; If the document is marked as "relevant", it goes directly to the next stage; If the document is "neutral", choose to keep it or continue to decompose the problem for optimization based on specific needs; For "irrelevant" documents, delete this search document; In the output phase, the system will return the retained documents and interpretation results; The self-reflection optimization module generates a preliminary answer based on the document and the question, and determines whether the preliminary answer is reasonable. If it is reasonable, the final answer is generated. If it is unreasonable, self-reflection optimization is performed; the self-reflection optimization module includes: inputting the original question, rewritten question and screened document into the large language model to obtain the preliminary generated answer; performing self-reflection check on the preliminary generated answer through the large language model, and the self-reflection check includes: answer consistency check, answer completeness check, answer logic check and answer semantic accuracy check.

Citation Information

Patent Citations

  • Open domain natural language reasoning question-answering system and method driven by large language model

    CN116932708A

  • Bank system question and answer method, device and equipment, medium and program product

    CN118260396A

  • Knowledge intensive question reasoning and generating method based on LLM

    CN118798367A

  • Vehicle behavior abnormal mode identification method for vehicle-mounted terminal

    CN119089371A