Natural language question and answer framework, method and device based on self-reflection

By introducing a self-reflection mechanism into the natural language question and answer framework, using the iterative evaluation and correction process of the thinking chain, persuader and answerer modules, the problem of insufficient ability of large language models in complex reasoning and knowledge utilization is solved, the model self-learning and optimization is achieved, and the task performance effect is improved.

CN120011491APending Publication Date: 2025-05-16SHENZHEN UNIV +1

Patent Information

Application Number
CN202411851002.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Large language models such as ChatGPT and PaLM are less capable of using complex inference and complex knowledge than humans. The existing technology has not yet explored the ability of LLMs to judge the correctness, analyze the reasoning paths and correct the paths.

Method used

It provides a natural language question-and-answer framework based on self-reflection, including the thinking chain module, the persuader module and the answerer module. It evaluates and corrects the model's inference path through iteratively to realize model self-learning and optimization.

Benefits of technology

Through the self-reflection framework, the model can learn knowledge more comprehensively and in-depth, improve the model's learning ability and knowledge application ability, and improve the model's performance on common sense questions and answers and other tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011491A_ABST
    Figure CN120011491A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, in particular to a natural language question and answer framework, method and device based on self-reflection, and can solve the problem that large language models LLMs (Language Models) such as ChatGPT and PaLM in the prior art show excellent performance in various language understanding and generation tasks to a certain extent, and the problem that the Language Models are difficult to understand and generate can be solved to a certain extent. However, the ability of the method in the aspects of complex reasoning and complex knowledge utilization is still lower than the human level. The framework comprises: a thinking chain module for receiving an input question, outputting a model thinking process according to the input question, and obtaining and outputting an answer at the end; the persuaser module firstly evaluates the correctness of the answer and the reasoning step, and outputs a corrected reasoning path to the responder module if an error exists in the reasoning step; and the answerer module is used for providing answers according to the corrected reasoning path and self-prompted question type information, and optimizing the output accuracy of the large model through repeated iteration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of large language models, and more specifically, to a natural language question-answering framework, method, and device based on self-reflection. Background Art

[0002] Hint learning has become an emerging paradigm in recent years. The idea of ​​hints originated from the experience of processing GPT. Researchers found that by carefully designing the input, such as choosing a suitable prefix of a sentence, the model can subsequently generate satisfactory content. This discovery inspired the study of similar properties of masked language models. Generally speaking, hints are intended to mimic the pre-training process of pre-trained language models (LMs). There are two main types of hints, namely prefix hints and cloze hints. The former, such as GPT, follows a left-to-right approach; the latter, such as BERT, follows the masked language modeling style and RoBERTa. The goal of the hint is to generate meaningful content, also known as the answer, at the masked position.

[0003] Recent advances in large-scale pre-trained language models (LLMs) have demonstrated impressive results in various natural language understanding and generation tasks. However, fully leveraging these models to solve complex problems, such as arithmetic problems, remains a challenging task. One cutting-edge technique, Chain-of-Thought (CoT) prompting, has achieved impressive results by forcing LLMs to solve problems step by step. Another technique, called In-Context Learning (ICL), is able to guide LLMs to generate the desired output by giving examples in context, demonstrating the effectiveness of prompt-based approaches in various applications. Based on these findings, existing techniques are dedicated to fully leveraging the output of the model and exploring capabilities beyond simple question answering, one of which is to judge and analyze the reasoning paths of the answers and correct them when necessary.

[0004] However, existing technologies only focus on using true answers or heuristic information as supervisory information to improve the reasoning path, but have not yet explored the ability of LLMs themselves to judge correctness, analyze reasoning paths, and correct paths. Summary of the invention

[0005] In order to solve the problem that large language models (LLMs) (Large Language Models) in the prior art, such as ChatGPT and PaLM, have shown excellent performance in various language understanding and generation tasks, but their capabilities in complex reasoning and complex knowledge utilization are still inferior to human levels, the present application provides a natural language question answering framework, method and device based on self-reflection.

[0006] The embodiment of the present application is implemented as follows:

[0007] In a first aspect, the present application provides a natural language question-answering framework based on self-reflection, including:

[0008] The thinking chain module receives input questions, outputs the model thinking process based on the input questions, and outputs the answer at the end;

[0009] The persuader module first evaluates the correctness of the answer and the reasoning steps. If there are errors in the reasoning steps, it outputs the corrected reasoning path to the respondent module;

[0010] The answerer module provides answers based on the corrected reasoning path and self-prompted question type information, optimizing the output accuracy of the large model through repeated iterations.

[0011] In one possible implementation, the framework is initialized by the thought chain module, which is represented by P. By providing input x and using model f, the system generates an answer and its corresponding reasoning path. This process can be formally represented as a function:

[0012] f: (x, P) → (r, a);

[0013] Among them, x is the question, r represents the sequence of tokens that constitute the reasoning path, and a represents the sequence of tokens containing the answer, usually at the end of the generated output.

[0014] In a possible implementation, the persuader module is represented by f C , taking x and (r, a) as input, the persuader module produces a modified reasoning path in which the last reasoning step is corrected, and its function can be defined as:

[0015] f C :(x,r,a)→(o crt , o aly , r c ).

[0016] In a possible implementation, the respondent module, denoted as fA, is a thought chain prompt that forces a single-step reasoning. Taking the reasoning path ^r from the persuader module as input, the respondent generates the supplementary information required to solve the given problem and provides an extended reasoning path, which includes an additional reasoning step, represented as a function:

[0017] f A :(x,^r)→(o typ , r a ).

[0018] In a possible implementation, the framework further includes iterating to obtain the completed reasoning path r from the answerer module. a After that, it is provided as input to the persuader module, thus starting a new iteration within the framework. The iterative process can be encapsulated by the combination of various functions, namely Through multiple modules, multiple loops are performed.

[0019] In one possible implementation, multiple answer choices are generated by the persuader module, converting open-ended questions into closed-ended questions by simply adding "wrong" after "correctness:", encouraging the model to generate diverse and reasonable answers.

[0020] In a second aspect, the present application provides a natural language question answering method based on self-reflection, including:

[0021] Input a question and output an initial answer based on the question;

[0022] Evaluate the correctness of initial answers and initial reasoning steps;

[0023] If the initial reasoning step is wrong, the initial reasoning step is modified and the corrected reasoning path is output.

[0024] Providing a revised answer and self-prompting question type information according to the revised reasoning path;

[0025] Output the revised answer, and re-evaluate the correctness of the revised answer and revised reasoning path for iteration.

[0026] In a third aspect, the present application provides a natural language question-answering device based on self-reflection, comprising:

[0027] The question input module is used to input questions and output initial answers based on the questions;

[0028] The answer evaluation module is used to evaluate the correctness of the initial answer and the initial reasoning steps;

[0029] A path correction module is used to modify the initial reasoning step if the initial reasoning step is wrong and output a corrected reasoning path;

[0030] An answer correction module, used for providing a corrected answer and self-prompted question type information according to the corrected reasoning path;

[0031] The answer output module is used to output the revised answer and re-evaluate the correctness of the revised answer and revised reasoning path for iteration.

[0032] The technical solution provided by this application can at least achieve the following beneficial effects:

[0033] The natural language question-answering method based on self-reflection provided in this application realizes the automatic internalization learning of language model knowledge by adopting a framework design in which three modules drive each other, making up for the deficiency of traditional externalized knowledge enhancement methods that rely on artificial knowledge. In particular, by using the persuader module, the model results can be questioned from multiple angles, driving the model to learn knowledge more comprehensively and deeply, strengthening the learning ability of the model. The application of the answerer module can prompt the model not only to answer questions, but more importantly, to learn how to use knowledge for logical expression, improving the model's ability to understand and apply knowledge. At the same time, the answer selection construction method is used to record the learning results of each module, realizing the refinement and accumulation of knowledge in the learning process, and improving the quality and consistency of the model output. The automated internalization learning framework is more convenient and flexible to implement and apply than the method that relies on external knowledge. The experiment proved that the framework can effectively improve the performance of the model in tasks such as common sense question answering. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0035] Figure 1 is a framework diagram of a natural language question-answering framework based on self-reflection shown in an exemplary embodiment of the present application;

[0036] Figure 2 is a schematic diagram of an iterative algorithm shown in an exemplary embodiment of the present application;

[0037] Figure 3 This is a schematic diagram of the framework of respondent priority reasoning shown in an exemplary embodiment of the present application.

[0038] Figure 4 is a flowchart of a natural language question answering method based on self-reflection shown in an exemplary embodiment of the present application;

[0039] Figure 5 It is a structural diagram of a natural language question-answering device based on self-reflection shown in an exemplary embodiment of the present application.

[0040] Reference numerals:

[0041] 1. Question input module; 2. Answer evaluation module; 3. Path correction module; 4. Answer correction module; 5. Answer output module. DETAILED DESCRIPTION

[0042] In order to make the purpose, implementation mode and advantages of the present application clearer, the exemplary implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, not all of the embodiments. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application.

[0043] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0044] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0045] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0046] Before explaining the natural language question-answering method based on self-reflection provided in the embodiment of the present application, the application scenario and implementation environment of the embodiment of the present application are first introduced.

[0047] Hint learning has become an emerging paradigm in recent years. The idea of ​​hints originates from the experience of dealing with GPT. Researchers have found that by carefully designing the input, such as choosing a suitable prefix of a sentence, the model can subsequently generate satisfying content. This discovery has motivated research on similar properties of masked language models. Generally speaking, hints are intended to mimic the pre-training process of a pre-trained language model (LM). There are two main types of hints, namely prefix hints and cloze hints. The former, such as GPT, follows a left-to-right approach; the latter, such as BERT, follows a masked language modeling style and RoBERTa. The goal of the hint is to generate meaningful content, also known as the answer, at the masked position. Consider a hint of the following form

[0048] [X]tpl([Z])

[0049] where slot [X] is usually filled with the input sentence, slot [Z] can be learned to be filled with some appropriate answer, and tpl is the template of the prompt. For example, consider the following prompt,

[0050] [Steve Jobs co-founded Apple with Steve Wozniak.]

[0051] Steve Jobs is the founder of Apple.

[0052] The prompt template tpl is usually designed manually, and the answer Z can be designed as a discrete choice from a given candidate set

[0053] Z = {employer, co-founder, manager, ...}

[0054] The answers to fill in the blanks can be in discrete form as above or continuous. In addition to the filled answers, other tokens constructed by the prompt can also be discrete words or learnable continuous vectors.

[0055] With the continuous development of prompt learning, it has gradually been applied to various task scenarios. For example, one of the major applications of prompts is as a parameter-efficient fine-tuning technique, which is used to fine-tune causal language models such as GPT. One of its major advantages is that by fine-tuning only a small number of parameters, the performance equivalent to full fine-tuning can be achieved. For example, P-tuning and P-tuning v2 are very representative works. It is precisely because of its small number of parameters and the ability to be easily decoupled from the ontology model that it is also widely used in customized tasks.

[0056] In the era of large models, many models with excellent performance have not been open sourced, such as ChatGPT and Claude. These large models are decoder-only GPT-like models, which means that prompts are very important for the performance of the model. Therefore, a large number of works related to "prompt engineering" have emerged. The main difficulty in this area is that because the model is closed source, it is impossible to directly search for the best prompt through gradient descent. Therefore, some works hope to filter out better prompts by constructing some non-differentiable supervision signals; and Zou et al. found that the prompts composed of discrete words they obtained by searching on the open source model have good transferability. Specifically, the discrete prompts they searched for on Llama-7B for the purpose of attack can also be used on ChatGPT to effectively attack it and generate harmful content. However, the automatic prompt construction of closed-source models is still very limited in nature, and the context learning and thinking chain that will be discussed next can make up for this well.

[0057] Contextual Learning

[0058] The increase in model size and corpus dimensionality has enabled large language models (LLMs) to demonstrate the ability to learn in context (ICL) - this feature allows large language models (LLMs) to complete the required tasks with only a small number of task-specific examples during reasoning, without changing the model's parameters. This provides an additional option for people to use complex cues that require search or optimization.

[0059] For example, for closed-source models, it is more feasible to use manually constructed contextual learning prompts than to use a single automatically constructed prompt. At the same time, the model can learn the user-defined goals from a small number of contextual prompt examples. As shown in Table 1, we want the model to perform sentiment analysis on a certain input. In addition to directly prompting it with "Please perform sentiment analysis on the following sentence", we can also prompt the model to output what we want by constructing examples (demonstration). Accordingly, how to construct and select appropriate examples has become a major focus of research in this field, and this module that can provide appropriate examples is also called an exampler.

[0060]

[0061] Table 1. Examples of contextual learning paradigms

[0062] This exploitation of demonstrations fosters a deeply intertwined relationship between Chain of Thought (CoT) prompts and ICL, and has therefore driven a great deal of research interest in identifying strategies that work well with few-example demonstrations. Since the introduction of few-example prompts, a number of methodologies have emerged that aim to enhance the prompting capabilities of models. These include automating prompt learning as well as providing task-specific instructions to models. Furthermore, the exploitation of demonstrations has fostered a deep connection between in-context learning (ICL) and Chain of Thought (CoT) prompts.

[0063] Thought chain tips

[0064] With the advancement of Large-Scale Language Models (LLMs) and In-Context Learning (ICL), especially the emergence of a series of novel prompt-related works, natural language models have made new breakthroughs in many natural language processing (NLP) downstream tasks involving reasoning and decision-making. Chain of Thoughts (CoT) prompt is a gradient-free technique that induces LLMs to generate intermediate reasoning steps leading to the final answer. Studies have shown that LLMs are able to perform CoT reasoning with zero-shot prompts or a few manually crafted example demonstrations.

[0065] The original idea of ​​the thought chain prompt technology is very simple. Researchers found that as long as they add "Let's think step by step" after the question when asking the model, they can stimulate the model's multi-step thinking ability (multi-step reasoning), thereby enabling the model to achieve better performance on the corresponding task.

[0066]

[0067] Table 2. Tips for constructing thought chains

[0068] Although mindchaining and related work have performed well in many traditional NLP reasoning and decision-making tasks, they still have limitations when facing complex logical reasoning or multi-hop problems.

[0069] ReAct explores the integration of LLMs in simultaneously generating reasoning traces and task-specific behaviors, promoting greater synergy between them. This approach builds a general framework that uses language models to combine reasoning and action to tackle a variety of language reasoning and decision-making tasks. The Describe, Explain, Plan, and Choose (DEPS) approach employs multi-step reasoning and subtask error correction to solve long-term tasks. Although DEPS shows significant performance by explaining errors in subtasks across trials, it relies on immediate failure detection of subtasks and fails to consider errors that may occur across a wide range of actions and subtasks.

[0070] Furthermore, DERA’s work is both fascinating and insightful. With GPT-4 being able to conduct robust and realistic conversations, the approach adopts dialogue as the medium for interaction. This approach frames the conversation as a discussion between two agent types – a researcher, who processes information and identifies the key components of a problem, and a decision maker, who has the autonomy to integrate the researcher’s information and make judgments on the final output.

[0071] Self-consistency adopts a synthetic process similar to sampling multiple times from the same thought chain prompt. It initially samples a diverse set of reasoning paths instead of relying solely on a greedy approach, and then selects the most consistent answer by marginalizing the sampled reasoning paths.

[0072] Large model self-enhancement

[0073] For large models, self-enhancement is an iterative enhancement based on prompts.

[0074] Different from self-consistency, there is a line of work focused on improving the output of LLMs using iterative methods. Self-Refine proposes a framework designed to enhance the initial output of LLMs through iterative feedback and refinement. The core concept revolves around using LLM to generate output, obtaining multi-faceted feedback about its own output from the same model, and then refining the previously generated output based on these feedbacks. It is worth noting that this iterative refinement framework does not require supervised training data or reinforcement learning, it only relies on a single LLM to operate.

[0075] Reflexion represents a significant improvement over previous methods such as ReAct and DEPS. By adopting a binary reward model, reflection enables agents to have dynamic memory and self-reflection capabilities, effectively enhancing their reasoning traces and task-specific behavior selection capabilities. At the same time, Iter-CoT adopts an iterative strategy to enhance reasoning steps and answers by leveraging correct answers as supervisory information, and finally generates typical examples by sampling from promoted examples. Similarly, PHP generates answers by leveraging hints from previous rounds until no new answers are generated.

[0076] Large Language Models (LLMs) such as ChatGPT and PaLM have demonstrated outstanding performance in various language understanding and generation tasks, but their capabilities in complex reasoning and sophisticated knowledge utilization are still inferior to human levels. Recent studies have identified the effectiveness of prompts in guiding LLMs to generate desired outputs.

[0077] Recent advances in large-scale pre-trained language models (LLMs) have demonstrated impressive results in various natural language understanding and generation tasks. However, fully leveraging these models to solve complex problems, such as arithmetic problems, remains a challenging task. A cutting-edge technique, the Chain-of-Thought (CoT) hint, achieves striking results by forcing LLMs to solve the problem step by step.

[0078] Another technique, called In-Context Learning (ICL), is able to guide LLMs to produce desired outputs by giving examples in context, demonstrating the effectiveness of prompt-based approaches in a variety of applications. Based on these findings, recent literature is dedicated to fully exploiting the output of the model and exploring capabilities beyond simple question answering. One of these capabilities is to judge and analyze the reasoning paths of the answers and correct them when necessary.

[0079] However, previous studies have only focused on using true answers or heuristic information as supervision information to refine the reasoning path, but have not explored the ability of LLMs to judge correctness, analyze reasoning paths, and correct paths by themselves. We call this ability "self-persuasion", that is, generating confident outputs based on given question-answer pairs and reasoning paths.

[0080] Based on this, the present application provides a natural language question-answering framework, method and device based on self-reflection, which utilizes the potential of large-scale pre-trained language models to iteratively improve the performance of LLMs. Some embodiment frameworks of the present application include three components: Normal CoT, a Convincer and an Answerer. It processes the output of a typical few example thinking chain prompts, evaluates the correctness of the response, reviews the answer, improves the reasoning, and finally generates a new solution.

[0081] Next, the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems will be described in detail through embodiments and in combination with the accompanying drawings. The embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Obviously, the described embodiments are part of the embodiments of the present application, not all of them.

[0082] In an exemplary embodiment, a natural language question answering framework based on self-reflection is provided. In this embodiment, the framework includes:

[0083] The thinking chain module receives the input question and outputs a question-answer pair based on the output question, and outputs the final answer output;

[0084] The persuader module first evaluates the correctness of the answer and the reasoning steps, and outputs the corrected reasoning path to the respondent module if the reasoning step produces an incorrect answer;

[0085] The respondent module provides answers based on the corrected reasoning path, as well as self-prompting question type information.

[0086] In one possible implementation, the framework includes a conventional CoT, a persuader, and a respondent, and the operation flow is as follows:

[0087] (1) Given a question-answer pair generated by any regular CoT (initialization);

[0088] (2) The persuader module first evaluates the correctness of the answer and reasoning steps (introspection);

[0089] (3) If the reasoning step produces an incorrect answer, the persuader module outputs the corrected reasoning path to the respondent module. The respondent then provides an answer based on the corrected reasoning path, as well as self-prompted question type information (answer);

[0090] (4) Finally, the regular CoT module completes the output of the last module (Completion).

[0091] Furthermore, by feeding the output of step (4) back to step (2), an iterative loop is formed.

[0092] Figure 1 It is a framework diagram of a natural language question-answering framework based on self-reflection shown in an exemplary embodiment of the present application.

[0093] In a possible implementation, the specific implementation is as follows:

[0094] The framework is named Self-Persuasion and consists of three basic components, namely, Conventional CoT, Persuaders, and Responders, which together facilitate four sequential steps, namely, “Initialize,” “Introspect,” “Answer,” and “Complete.”

[0095] like Figure 1 As shown, the left part shows an "iterative loop"; the middle and right parts show the details of the input and output of each module. The prompts of each module are highlighted in "Input", and the next stage output of each module is also highlighted in "Output". Balloon marks indicate where the respondent's priority reasoning occurs. Asterisks mark the output to the next iteration.

[0096] In the initial phase, the regular CoT generates an initial answer to the input question, called the “regular output”, which is then fed as input to the persuader. Detailed insights into the functioning of the persuader can be found in Figure 1 , where it generates three key outputs based on the input question and the "Regular Output": "Correctness", "Analysis", and "Final Answer". If the model considers the answer to be correct, as indicated by Correctness: Correct in the "Persuader Output", then the original answer in the "Regular Output" is retained as the final answer and the "Introspection" step is repeated.

[0097] Conversely, if the persuader determines that the answer is wrong, the "final answer" from the "persuader output" is fed into the last step, "answer". This step uses the answerer to generate an intermediate answer. Finally, the regular CoT is used to "complete" the answer. Subsequently, the resulting answer is fed into the "introspection" step as the "regular output" of the new iteration.

[0098] in:

[0099] General thought chain tips

[0100] The self-persuasion framework is initialized by using an arbitrary thought chain prompt denoted as P. By providing an input x and leveraging the model f, the system generates an answer and its corresponding reasoning path.

[0101] This process can be formally represented as a function f:(x,P)→(r,a), where r represents the sequence of tokens that make up the reasoning path and a represents the sequence of tokens that contains the answer, usually at the end of the generated output.

[0102] To simplify the representation, we can represent the model with hint P as f P , and the revised function expression is obtained: f P : x→(r,a). In subsequent chapters, we will use similar notations for clarity and consistency.

[0103] Marginal Persuaders

[0104] The second key component of the self-persuasion framework is the persuader, denoted as f C , which plays a core role in the system. Taking x and (r, a) as input, the persuader module produces a modified reasoning path in which the last reasoning step is corrected. Its function can be defined as f C :(x,r,a)→(o crt , o aly , r c ). The persuader has the ability to identify errors in the original reasoning path r, analyze the errors, and then correct them. Once the erroneous reasoning step is corrected, the subsequent steps will be removed from the reasoning path.

[0105] like Figure 1 As shown, the persuader module can detect errors in the reasoning path, especially in the paragraph "...So, 2 / 3*x-10=40+1 / 3*x. Simplifying this equation, we get 1 / 3*x=90...". After analyzing the error, it corrects the error and outputs a truncated answer, such as "Let the number be x. So, 2 / 3*x-10=40+1 / 3*x. Simplifying this equation, we get 1 / 3*x=50, which means x=150." It should be noted that a reasoning path may contain multiple errors. Therefore, in the design of the persuader module, we focus on analyzing and correcting the first identified error while marginalizing the remaining steps.

[0106] Step by step answerer

[0107] The third building block of our framework is called the “Answerer”, denoted as fA. It can be viewed as a thought chain prompt that forces a single-step reasoning. Taking the reasoning path ^r from the persuader module as input, the answerer produces the supplementary information required to solve the given question. In addition, it provides an extended reasoning path that includes an additional reasoning step. This can be formally represented as the function f A :(x,^r)→(o typ , r a ).

[0108] like Figure 1 As shown in the running example in , the respondent module appends relevant information to the answer of the previous step, such as "Type: Algebraic equation". By adopting this step-by-step approach, the respondent gradually contributes to the reasoning process and thus gives a comprehensive and powerful answer to the question asked.

[0109] Figure 2 It is a schematic diagram of an iterative algorithm shown in an exemplary embodiment of the present application.

[0110] Iteration

[0111] like Figure 2 As shown, the completed reasoning path r is obtained from the respondent module a After that, we can provide it as input to the persuader module, thus starting a new iteration within the framework. This iteration process can be encapsulated by the combination of various functions, namely By utilizing these building blocks, the algorithm performs multiple loops, such as Figure 2 shown.

[0112] Answer Selection Build

[0113] We observed that in question-answering tasks without answer choices, the persuader module tends to produce responses like “Correctness: True.”. To address this limitation, we leverage the persuader module to generate multiple answer choices, effectively turning open-ended questions into closed-ended questions. This is achieved by simply adding “False.” after “Correctness:”, thereby encouraging the model to generate a diverse set of plausible answers. We can think of this as creating a “what-if world” where correctness is a given and as a premise may be false in certain situations. For simplicity, this module is named the Error-Only Persuader.

[0114] Figure 3 It is a schematic diagram of the framework of respondent priority reasoning shown in an exemplary embodiment of the present application.

[0115] Answer priority

[0116] like Figure 3As shown above, the persuader module not only provides the correctness evaluation of the initially generated answer, but also crt , and also provides a revised reasoning path r c However, expecting LLMs to modify their reasoning paths can pose challenges. In some cases, the persuader module may accurately evaluate correctness but have difficulty in correcting the reasoning path accordingly. To address this issue, we propose an alternative reasoning approach, called answerer-first reasoning, which avoids relying on correcting the reasoning path r c This is done by passing the output of the answerer module o typ Incorporation into conventional CoT f P This is done by appending o after "A:". typ ,like Figure 3 By adopting this approach, we enable the model to prioritize the answerer module during inference.

[0117] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the instructions, these steps are not necessarily executed in sequence according to the order of the instructions. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.

[0118] Corresponding to the aforementioned embodiment of a natural language question-answering framework based on self-reflection, adopting the same technical concept, the present application also provides an embodiment of a natural language question-answering method based on self-reflection.

[0119] Figure 4 It is a flowchart of a natural language question-answering method based on self-reflection shown in an exemplary embodiment of the present application.

[0120] In an exemplary embodiment, Figure 4 As shown, a natural language question answering method based on self-reflection is provided. In this embodiment, the method may include the following steps:

[0121] Step 100: Input a question and output an initial answer based on the question.

[0122] Step 200: Evaluate the correctness of the initial answer and initial reasoning steps.

[0123] Step 300: If the initial reasoning step is wrong, the initial reasoning step is modified and a corrected reasoning path is output.

[0124] Step 400: Providing a revised answer and self-prompting question type information according to the revised reasoning path.

[0125] Step 500: Output the revised answer, and re-evaluate the correctness of the revised answer and the revised reasoning path, and perform iteration.

[0126] Corresponding to the aforementioned embodiment of a natural language question-answering method based on self-reflection and adopting the same technical concept, the present application also provides an embodiment of a natural language question-answering device based on self-reflection.

[0127] Figure 5 It is a structural diagram of a natural language question-answering device based on self-reflection shown in an exemplary embodiment of the present application.

[0128] In an exemplary embodiment, Figure 5 As shown, the natural language question-answering device based on self-reflection includes:

[0129] The question input module 1 is used to input questions and output initial answers according to the questions.

[0130] The answer evaluation module 2 is used to evaluate the correctness of the initial answer and the initial reasoning steps.

[0131] The path correction module 3 is used to modify the initial reasoning step if the initial reasoning step is wrong and output a corrected reasoning path.

[0132] The answer correction module 4 is used to provide a corrected answer and self-prompted question type information according to the corrected reasoning path.

[0133] The answer output module 5 is used to output the revised answer and re-evaluate the correctness of the revised answer and the revised reasoning path for iteration.

[0134] For the specific limitations of the natural language question-answering device based on self-reflection, please refer to the limitations of the natural language question-answering method based on self-reflection above, which will not be repeated here. Each module in the above-mentioned natural language question-answering device based on self-reflection can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0135] It can be seen that some embodiments of the present application propose the use of three types of modules: a conventional thought chain prompt module, a marginalized persuader module, and a step-by-step answerer module. These three modules play different roles, such as prompt learning, questioning learning, and in-depth learning. These three modules are combined together to form a framework that can realize self-learning and optimization of the model. The key is to use the interactive and collaborative relationship of the three modules to drive the model to continuously improve and learn, and ultimately achieve self-enhancement of knowledge. This method conception innovatively internalizes the learning process of the language model, thereby realizing automatic enhancement of knowledge.

[0136] It also proposes the use of a marginal persuader module, which can question the results given by the model. Questioning includes checking the marginal details of the results, consistency, etc., and checking whether the answer is completely correct from different angles. Through questioning, the model can reflect on whether there are problems with the answer and make suggestions for modification. This module uses questioning to drive the model to self-correct and correct the framework design of the original answer.

[0137] It also proposes the use of a step-by-step answerer module, which can answer questions in a way of further explanation and argumentation. This module will require the model to provide a more in-depth and comprehensive discussion of the original answer, and give reasons or basis. This is different from a direct answer, and strengthens the model's ability to use knowledge for argumentation and expression. Driven by this module, the model can continuously improve and expand its understanding and expression of the problem. Combined with other modules, it can help the model learn related knowledge while answering a question.

[0138] In the answer choice construction method, we use the persuader module to generate multiple answer choices, effectively turning open-ended questions into closed-ended questions. This is achieved by simply adding "False." after "Correctness:", thereby encouraging the model to generate diverse and reasonable answers.

[0139] There is also respondent-first theory to avoid relying on the modified reasoning path r c , the output o of the respondent module typ Incorporation into conventional CoT f P ,By adopting this approach, we make the model prioritize the answerer module during ,inference.

[0140] In order to verify the feasibility of this application:

[0141] In the evaluation of some embodiments of the present application, we apply the self-persuasion framework to a series of different benchmarks, including the following datasets: (1) GSM8K, a benchmark focusing on math word problems, (2) AddSub, a collection of addition and subtraction problems, (3) SVAMP, a dataset containing challenging math word problems, (4) AQuA, a dataset containing algebra word problems, (5) CSQA, a dataset containing common sense questions, (6) Date Understanding, a common sense dataset focusing on date reasoning problems, (7) and GAOKAO, a benchmark based on the Chinese college entrance examination covering a variety of question types. For some embodiments of the present application, we consistently use temperature sampling in all results, with a temperature value of T = 0.7.

[0142] We test on the development set of CSQA while restricting the evaluation to math problems on the GAOKAO benchmark. In particular, the specific prompts we use are provided in the Appendix. For the GAOKAO benchmark, we adopt a one-shot persuader approach combined with a zero-shot answerer approach. Unless explicitly stated otherwise, all ablation experiments are conducted on AQuA. The dataset details are detailed in Table 3.

[0143]

[0144] Table 3. Benchmark datasets for the experiments. Since this work is zero-shot or few-shot learning, some of the selected datasets do not have any training data.

[0145] Results: The main results of self-persuasion and comparison with other methods are shown in Table 4. We report the results of the last iteration (up to 5 iterations), and the scores of each iteration are shown in Tables 5 and 6. The results of each iteration step of the two reasoning types are shown. Some embodiment methods of the present application start with the results of ordinary CoT, marked as iteration 0. We considered several previous works for comparison. The table includes two methods using "GPT-turbo-3.5" and their results using "ordinary CoT". Since there is a large gap between the initial results of the listed methods, we also put the average and improvement of the results in the table. It is worth mentioning that although our initial results on AQuA are lower than PHP, self-persuasion still surpasses their final results by a large margin.

[0146]

[0147] Table 4. Main results on different datasets. The best results are shown in bold.

[0148] .Effectiveness of Self-Persuasion on Challenging Arithmetic Tasks: When evaluating English arithmetic tasks, we can rank them by difficulty as follows: “AQuA > GSM8K ≈ SVAMP > AddSub”, and Table 4 shows that Self-Persuasion exhibits significant improvements on all types of reasoning. In particular, it achieves the highest improvement on AQuA (+6.1% on average), followed by GSM8K (+4.1%), SVAMP (+4.0%), and AddSub (+0.8%).

[0149]

[0150]

[0151] Table 5. Average accuracy and standard deviation of each step. We run twice to calculate the results.

[0152]

[0153] Table 6 Results of only wrong persuaders at different iteration numbers

[0154] Diversity and potential of self-persuasion in different problem domains: Although a slight drop is observed in AddSub, which is consistent with the findings of PHP, our proposed method shows significant performance improvement compared to the initial plain CoT method in all other benchmarks. It is worth noting that some of the embodiment methods of this application perform well in arithmetic tasks, surpassing PHP even when the initial performance of plain CoT is lower than PHP in SVAMP and AQuA. In addition, the results obtained on the common sense dataset provide strong evidence for the potential applicability of some embodiment frameworks of this application in a wider range of problem domains.

[0155] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0156] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A natural language question answering framework based on self-reflection, characterized in that: include: The thinking chain module receives input questions, outputs the model thinking process based on the input questions, and outputs the answer at the end; The persuader module first evaluates the correctness of the answer and the reasoning steps. If there are errors in the reasoning steps, it outputs the corrected reasoning path to the respondent module; The answerer module provides answers based on the corrected reasoning path and self-prompted question type information, optimizing the output accuracy of the large model through repeated iterations.

2. The self-reflection-based natural language question-answering framework according to claim 1, characterized in that: The framework is initialized by the thought chain module, which is denoted as P. By providing input x and using the model f, the system generates an answer and its corresponding reasoning path. This process can be formally expressed as a function: f: (x, P) → (r, a); Among them, x is the question, r represents the sequence of tokens that constitute the reasoning path, and a represents the sequence of tokens containing the answer, usually at the end of the generated output.

3. The self-reflection-based natural language question-answering framework according to claim 2, characterized in that: The persuader module is denoted as f C , taking x and (r, a) as input, the persuader module produces a modified reasoning path in which the last reasoning step is corrected, and its function can be defined as: f C :(x,r,a)→(o crt ,o aly ,r c )。 4. The self-reflection-based natural language question-answering framework according to claim 3, characterized in that: The respondent module, denoted as f A , is a thought chain prompt that forces a single-step reasoning. With the reasoning path ^r from the persuader module as input, the respondent generates the supplementary information needed to solve the given problem and provides an extended reasoning path that includes an additional reasoning step, expressed as a function: f A :(x,^r)→(o typ ,r a )。 5. The self-reflection-based natural language question-answering framework according to claim 4, characterized in that: The framework also includes iterating to obtain the completed reasoning path r from the answerer module a After that, it is provided as input to the persuader module, thus starting a new iteration within the framework. The iterative process can be encapsulated by the combination of various functions, namely Through multiple modules, multiple loops are performed.

6. The self-reflection-based natural language question-answering framework according to claim 5, characterized in that: Multiple answer choices are generated through the persuader module, turning open-ended questions into closed-ended questions by simply adding "false" after "correctness:", encouraging the model to generate diverse and reasonable answers.

7. A natural language question answering method based on self-reflection, characterized in that: include: Input a question and output an initial answer based on the question; Evaluate the correctness of initial answers and initial reasoning steps; If the initial reasoning step is wrong, the initial reasoning step is modified and the corrected reasoning path is output. Providing a revised answer and self-prompting question type information according to the revised reasoning path; Output the revised answer, and re-evaluate the correctness of the revised answer and revised reasoning path for iteration.

8. A natural language question-answering device based on self-reflection, characterized in that: include: The question input module is used to input questions and output initial answers based on the questions; The answer evaluation module is used to evaluate the correctness of the initial answer and the initial reasoning steps; A path correction module is used to modify the initial reasoning step if the initial reasoning step is wrong and output a corrected reasoning path; An answer correction module, used for providing a corrected answer and self-prompted question type information according to the corrected reasoning path; The answer output module is used to output the revised answer and re-evaluate the correctness of the revised answer and revised reasoning path for iteration.

Citation Information

Patent Citations

  • Large model training method and system based on self-provincial inversion

    CN117851829A

  • Method and device for generating answers

    CN118396108A

  • Dual-model reflection learning method based on thinking chain correction and adaptive screening

    CN118747530A

Cited By

  • Automatic thinking chain prompt generation method based on black box optimization and vulnerability quantification

    CN120336491A

  • Training data synthesis method and device based on error extrapolation and inference chain analysis, medium and program product

    CN120611192A

  • Training data synthesis method and device based on error extrapolation and reasoning chain analysis, medium and program product

    CN120611192B