Question and answer method and device based on large language model, equipment and medium

By assigning weights and information entropy to each reasoning step of a large language model, constructing relative advantage values, and optimizing the model to improve the accuracy of key steps, the problems of logical jumps and factual errors in large language models in complex problems are solved, achieving higher accuracy of reasoning processes and reliability of answers.

CN121920574AActive Publication Date: 2026-04-24NEW H3C TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NEW H3C TECH CO LTD
Filing Date
2026-03-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing large language models suffer from logical jumps, factual errors, and insufficient evidence in the multi-step reasoning process of complex problems, resulting in a final answer that is formally reasonable but substantively incorrect, and lacking assurance of the accuracy of intermediate reasoning steps.

Method used

By assigning weights and information entropy to each reasoning step, a relative advantage value is constructed to optimize the large language model, thereby improving the accuracy and logical rigor of key steps, suppressing uncertainty interference, and guiding the model towards a more reliable reasoning process.

Benefits of technology

It significantly improves the accuracy of reasoning processes and the reliability of final answers in question-and-answer scenarios for large language models, especially in complex scenarios such as mathematical proofs, code generation, and scientific question answering, ensuring the accuracy and reliability of the answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920574A_ABST
    Figure CN121920574A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a question answering method and device based on a large language model, equipment and a medium. In the application, for each reasoning step of each answer, a reward value reflecting the quality of the reasoning step is set for the reasoning step, a weight quantifying the relative influence degree of the reasoning step on the quality of the answer is set for the reasoning step, and an information entropy representing the uncertainty degree of the reward value of the reasoning step is calculated for the reasoning step. On the basis, the reward value is cooperatively adjusted through the weight and the information entropy, and a relative advantage value capable of accurately measuring the relative value of the overall reasoning process of the answer is obtained. Relative advantage values of the M answers are used as optimization signals, and the large language model is guided to be optimized in the direction of focusing on a reasoning path for improving the accuracy of the key reasoning step and inhibiting uncertainty interference, so that the accuracy of the reasoning process of the large language model in a question and answer scene is improved, and the accuracy and reliability of generated answers are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to question-answering methods, apparatus, devices and media based on large language models. Background Technology

[0002] With the continuous breakthroughs in artificial intelligence technology, large language models (LLMs) are being widely used in various user question-and-answer scenarios.

[0003] In practical applications, large language models rely on multi-step reasoning to generate answers when dealing with complex problems such as mathematical reasoning, causal analysis, and programming. However, during the reasoning process, intermediate steps are prone to logical jumps, factual errors, or insufficient evidence, leading to a final answer that may be formally reasonable but substantively flawed.

[0004] Therefore, there is an urgent need for a new question-answering method based on a large language model to improve the accuracy of the reasoning process in question-answering scenarios and ensure the accuracy and reliability of the generated answers. Summary of the Invention

[0005] In view of this, embodiments of this application provide methods, apparatus, devices and media for improving question answering based on large language models, so as to improve the accuracy of the reasoning process in question answering scenarios and ensure the accuracy and reliability of the generated answers.

[0006] This application provides a question-answering method based on a large language model, the method comprising: For each answer output by the large language model, the weight of each reasoning step in the answer is determined; the weight of any reasoning step in the answer is used to represent the relative importance of the reasoning step to the correctness of the answer among all reasoning steps in the answer; the larger the weight of the reasoning step, the greater the relative importance of the reasoning step; the large language model outputs M answers based on the same input question; The information entropy of each reasoning step is determined by the reward value of each reasoning step in the answer. The information entropy of any reasoning step is used to represent the confidence level of the reward value of that reasoning step. The smaller the information entropy of any reasoning step, the greater the confidence level of the reward value of that reasoning step. Based on the reward value, weight, and information entropy of each reasoning step in each answer, the relative advantage value of the answer is determined. The relative advantage value of any answer is used to represent the relative value of the answer among M answers. The large language model is optimized based on the relative advantage value of each answer.

[0007] This application also provides a question-answering device based on a large language model, the device comprising: The first determining module is used to determine the weight of each reasoning step in each answer output by the large language model. The weight of any reasoning step in the answer is used to characterize the relative importance of the reasoning step to the correctness of the answer among all reasoning steps in the answer. The larger the weight of the reasoning step, the greater the relative importance of the reasoning step. The large language model outputs M answers based on the same input question, where M is greater than 1. The second determining module is used to determine the information entropy of each reasoning step in the answer by using the reward value of each reasoning step; the information entropy of any reasoning step is used to represent the confidence level of the reward value of the reasoning step; the smaller the information entropy of any reasoning step, the greater the confidence level of the reward value of the reasoning step. The third determining module is used to determine the relative advantage value of each answer based on the reward value, weight, and information entropy of each reasoning step in each answer. The relative advantage value of any answer is used to represent the relative value of the answer among M answers. An optimization module is used to optimize the large language model based on the relative advantage value of each answer.

[0008] This application also provides an electronic device, including: a processor and a machine-readable storage medium for storing machine-executable instructions, wherein the machine-executable instructions, when run by the machine-readable storage medium, cause the processor to perform the steps of the above method.

[0009] This application also provides a machine-readable storage medium storing machine-executable instructions that, when executed, enable the implementation of the steps described above.

[0010] As can be seen from the above technical solution, in this embodiment, firstly, a reward value reflecting the quality of each reasoning step is generated for each of the N reasoning steps in the question-and-answer sequence. Then, based on the importance of each reasoning step in the sequence, a corresponding weight is assigned, and an information entropy representing the uncertainty of the reward value for that reasoning step is calculated. Finally, by fusing the reward value, weight, and information entropy, a relative advantage value that accurately measures the relative value of the overall reasoning process of each answer is constructed. Using the relative advantage value of each answer as an optimization signal, the large language model is guided towards optimizing towards a more logically rigorous and coherent reasoning process, while suppressing low-quality or uncertain reasoning processes. This significantly improves the accuracy of the large language model's reasoning process in question-and-answer scenarios, ensuring the accuracy and reliability of the final generated answer.

[0011] Furthermore, by employing the method provided in the embodiments of this application, even in complex question-and-answer scenarios that highly rely on multi-step rigorous reasoning, such as mathematical proofs, code generation, and scientific question answering, the accuracy of the reasoning process can still be guaranteed, ensuring the accuracy and reliability of the final generated answer. Attached Figure Description

[0012] Figure 1 A flowchart illustrating the method provided in the embodiments of this application; Figure 2 A schematic diagram of the process for determining the relative advantage value of the answer provided in an embodiment of this application; Figure 3 Another flowchart illustrating the method provided in this application embodiment; Figure 4 This is a schematic diagram of the device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0014] Before introducing the method provided in the embodiments of this application, the existing technical problems will be explained: In various user question-and-answer scenarios, such as in a math question-and-answer scenario, a user inputs a math question, and the large language model generates an answer that can include the steps to solve the problem and the final correct answer. Alternatively, in a causal analysis user question-and-answer scenario, a user inputs a causal analysis question related to legal consultation, such as, "I slipped and fell on a sidewalk on a rainy day because of a banana peel that hadn't been cleaned up in time. Can I claim compensation from the municipal department?" The large language model constructs a causal chain based on this question and provides a liability determination.

[0015] In the aforementioned user question-and-answer scenarios, large language models often need to deduce the final answer through multiple steps of reasoning. However, existing large language models generally suffer from insufficient accuracy in intermediate reasoning steps, such as logical jumps, insufficient evidence, or factual errors (i.e., insufficient accuracy in the reasoning process), leading to a situation where the final answer may be formally reasonable but substantively incorrect.

[0016] In related technologies, Group Relative Policy Optimization (GPRO) is commonly used to optimize large language models. GPRO is a reinforcement learning algorithm for optimizing the training of large language models, aiming to reduce training resource consumption while maintaining the stability of parameter updates. When using GPRO to reinforce the training of a large language model, multiple answers are provided for a single input, and reward values ​​are assigned to each answer. The relative advantage of each answer is calculated based on its reward values, and the relative advantage of each answer is used to optimize the large language model, thereby increasing the probability of the large language model generating correct answers.

[0017] However, large language models optimized in this way usually focus on the correctness of the final answer and still lack guidance and calibration for intermediate reasoning steps. When faced with complex problems that require multiple reasoning steps, large language models still suffer from insufficient accuracy in the reasoning process.

[0018] Therefore, improving the accuracy of the reasoning process in question-and-answer scenarios to ensure the accuracy of the generated answers remains an urgent problem to be solved.

[0019] Based on this, embodiments of this application provide a question-answering method, apparatus, device, and medium based on a large language model to improve the accuracy of the reasoning process in question-answering scenarios and ensure the accuracy and reliability of the generated answers.

[0020] The method provided in the embodiments of this application is described in detail below: See Figure 1 , Figure 1 This is a flowchart illustrating the method provided in an embodiment of this application.

[0021] S101, For each answer output by the large language model, determine the weight of each reasoning step in the answer; the weight of any reasoning step in the answer is used to represent the relative influence of the reasoning step on the quality of the answer among all reasoning steps in the answer; the larger the weight of the reasoning step, the greater the relative influence of the reasoning step on the quality of the answer; the large language model outputs M answers based on the same input question; M is greater than 1.

[0022] In this step, the weight of any reasoning step in any answer is used to represent the relative influence (or relative contribution) of the quality of that reasoning step on the overall quality of the answer compared to other reasoning steps in that answer.

[0023] The greater the weight of any reasoning step in any answer, the more significant the impact of the correctness of that reasoning step on the correctness of the final answer, and the more significant the impact of its error on the error of the final answer, compared to other reasoning steps in that answer. Conversely, the smaller the weight of a reasoning step, the weaker the impact of its correctness or error on the correctness or error of the final answer, compared to other reasoning steps.

[0024] The weight of any reasoning step in any answer is predetermined based on its position within the total number of reasoning steps in that answer. This is because, generally, the quality of a reasoning step further down the line in the answer has a greater impact on the overall answer quality. Therefore, when assigning weights, later reasoning steps are given larger weights to reflect the principle that the quality of later reasoning steps has a greater relative impact on the final answer quality.

[0025] Optionally, as an embodiment, the weight of each reasoning step in any answer can be determined by: determining the linear weight of the reasoning step based on the step number of the reasoning step in the answer and the total number of reasoning steps in the answer.

[0026] For example, suppose the answer to each question generates N reasoning steps, and any reasoning step for any answer i... j The linear weights can be expressed by the following formula (1): Formula (1) in, The reasoning steps to answer i j Linear weights; I ij For reasoning steps j The step number in answer i; i∈[1,2,3..M]; M is the total number of answers to the same question; j ∈[1,2,3..N], where N is the total number of reasoning steps in answering i.

[0027] In this step, by assigning corresponding weights to each reasoning step of each answer, the large language model is guided to focus on the key reasoning steps that have a greater impact on generating the final answer during the optimization process. This guides the large language model to pay more attention to the more critical reasoning steps when optimizing, thereby guiding the large language model to optimize in the direction of improving the accuracy of key reasoning steps.

[0028] S102, using the reward value of each reasoning step in the answer, determine the information entropy of that reasoning step; the information entropy of any reasoning step is used to indicate the confidence level of the reward value of that reasoning step; there is a negative correlation between the information entropy of any reasoning step and the confidence level of the reward value of that reasoning step.

[0029] In related technologies, only a reward value is provided for a large language model to output M answers based on the same input question. The embodiments of this application improve upon this by providing a reward value for each reasoning step of each answer in the M answers.

[0030] The reward value of any reasoning step for any answer is used to evaluate the quality of that reasoning step, specifically to quantify the correctness of the reasoning step itself and its contribution to driving the entire reasoning process toward generating the correct final answer.

[0031] The higher the reward value of any reasoning step, the more accurate, well-founded, and logically consistent the reasoning step is, and the more crucial and positive its role in leading to the correct answer (i.e., the higher its contribution). Conversely, the lower the reward value of a reasoning step, the less accurate, unreasonable, and logically inconsistent its reasoning step is, and the lower its contribution in leading to the correct answer.

[0032] Specifically, there are many ways to provide a reward value for each reasoning step of each answer. For example, as an embodiment, the question, the answer to the question, each reasoning step of the answer, and all reasoning steps of the answer are input into the Process-supervised Reward Model (PRM). The PRM model evaluates the logical correctness of the reasoning step in all reasoning steps, its coherence with the preceding and subsequent reasoning steps, and its contribution to driving the entire reasoning process toward the final generation of the correct answer, and outputs the reward value of the reasoning step.

[0033] For example, the question is: Jane's duck lays 16 eggs a day. She eats three eggs every morning and uses four eggs to bake muffins for her friends. She sells the remaining eggs at the farmers market every day for $2 each fresh duck egg. How much money does she earn at the farmers market every day? The large language model's answer to this question, and the rewards provided by the PRM model for each reasoning step of this answer, are shown below: Reasoning step 1: Janet's ducks lay 16 eggs a day, and the reward for reasoning step 1 is [0.97].

[0034] Reasoning step 2: She eats three eggs every morning, so she has 16-3=13 eggs left. The reward for reasoning step 2 is [0.93].

[0035] Reasoning step 3: She bakes muffins for her friends with four eggs every day, so she has 13-4=9 eggs left. The reward for reasoning step 3 is [0.94].

[0036] Reasoning step 4: She sells the remaining eggs at the farmers market every day for $2 each fresh duck egg, so she can earn 9*2=$18 per day. The reward for reasoning step 4 is [0.99].

[0037] In this step, compared to related technologies that only generate a single reward value for the final answer, this application embodiment generates a reward value for each reasoning step. This helps guide large language models to learn more reliable and accurate multi-step reasoning capabilities, significantly improving the accuracy of intermediate steps and the final answer in complex problems.

[0038] Furthermore, considering that the uncertainty of the reward value in the reasoning step itself can interfere with the optimization of the large language model, if the reward value is directly used for question answering based on the large language model, it will interfere with the optimization of the large language model towards focusing on key reasoning steps and improving the accuracy of reasoning steps. Therefore, this application introduces information entropy to quantify the degree of uncertainty of the reward value, thereby suppressing the interference of the uncertainty of the reward value on question answering based on the large language model.

[0039] The information entropy of any reasoning step reflects the degree of uncertainty in evaluating the quality of that step. The degree of uncertainty is the opposite of the confidence level. Therefore, the information entropy of any reasoning step is negatively correlated with the confidence level of that reasoning step.

[0040] The smaller the information entropy of any inference step, the lower the uncertainty in predicting the reward value of that inference step, and the higher the confidence level of the reward value of that inference step. Conversely, the larger the information entropy of any inference step, the greater the uncertainty in predicting the reward value of that inference step, and the lower the confidence level of the reward value of that inference step.

[0041] The information entropy of any inference function can be calculated as follows: for each inference step, a specified operation is performed on the reward value of the inference step and the difference between the specified value and the reward value of the inference step, and the result of the specified operation is determined as the information entropy of the inference step.

[0042] For example, the information entropy of any reasoning step is calculated using the following formula (2): Formula (2) in, To answer the reasoning steps in ij The reward value; H j To answer the reasoning steps in i j Information entropy.

[0043] Optionally, as an embodiment, after obtaining the initial information entropy of N reasoning steps for each answer using the above formula, the average information entropy of the initial information entropy of each reasoning step is obtained, and the average information entropy is used as the final information entropy of each reasoning step for that answer.

[0044] In this step, by setting information entropy for each inference step to quantify the evaluation uncertainty of its reward value, the interference of low-confidence reward signals is suppressed in the question-answering process based on the large language model. This guides the large language model to pay more attention to inference steps with high confidence and high reliability, avoids incorrect optimization direction caused by noisy rewards, and effectively improves the accuracy of multiple inference steps.

[0045] S103. Based on the reward value, weight, and information entropy of each reasoning step in each answer, determine the relative advantage value of the answer. The relative advantage value of any answer is used to represent the relative value of the answer among the M answers.

[0046] In this step, weights reflect the relative impact of reasoning steps on the quality of the final answer, and information entropy reflects the uncertainty of reward evaluation. By adjusting the reward value through weights and information entropy, the obtained relative advantage value more accurately reflects "which key reasoning steps contributed to the relative value of the answer with what confidence level." This provides a clear optimization direction for subsequent optimization of the large language model: it can focus on improving the accuracy of more critical reasoning steps while suppressing the interference of reward uncertainty, thereby improving the accuracy of the reasoning process of the large language model in answering questions, and thus enhancing the accuracy and robustness of answering complex questions in question-answering scenarios.

[0047] As for how to determine the relative advantage value of an answer based on the reward value, weight, and information entropy of each reasoning step in each answer, this will be explained in detail later with specific examples, and will not be elaborated here.

[0048] S104, optimize the large language model based on the relative advantage value of each answer.

[0049] In this step, after obtaining the relative advantage values ​​of M answers, a loss function is constructed based on the relative advantage values ​​of the M answers for each of the multiple questions. The policy parameters of the large language model are updated according to the loss function until the stopping optimization condition is met. The large language model that meets the stopping optimization condition is determined as the target large language model.

[0050] This concludes the process. Figure 1 The process is shown below.

[0051] pass Figure 1 As shown in the flowchart, in this embodiment, firstly, a reward value reflecting the quality of each reasoning step is generated for each of the N reasoning steps in the answer. Then, based on the importance of each reasoning step in the reasoning step sequence, a corresponding weight is assigned, and the information entropy representing the uncertainty of the reward value for that reasoning step is calculated. Finally, by fusing the reward value, weight, and information entropy, a relative advantage value that accurately measures the relative value of the overall reasoning process of each answer is constructed. Using the relative advantage value of each answer as an optimization signal, the large language model is guided towards optimizing towards a more logically rigorous and coherent reasoning process, while suppressing low-quality or uncertain reasoning processes. This significantly improves the accuracy of the large language model's reasoning process in question-and-answer scenarios, ensuring the accuracy and reliability of the final generated answer.

[0052] Furthermore, by employing the method provided in the embodiments of this application, even in complex question-and-answer scenarios that highly rely on multi-step rigorous reasoning, such as mathematical proofs, code generation, and scientific question answering, the accuracy of the reasoning process can still be guaranteed, ensuring the accuracy and reliability of the final generated answer.

[0053] The following is a detailed explanation of how the relative advantage value of this answer was determined: See Figure 2 , Figure 2 This is a schematic diagram of the process for determining the relative advantage value of the answer provided in an embodiment of this application.

[0054] like Figure 2 As shown, the process may include the following steps: S201, perform a first processing on the first reward value sequence corresponding to the M answers to obtain a processing result; the first processing is used to make each reward value in the processing result meet a first set minimum difference requirement; wherein, the first reward value sequence corresponding to any answer is obtained by performing a second processing on the second reward value sequence corresponding to the answer; the second reward value sequence corresponding to any answer is a sequence formed by the reward values ​​of each reasoning step in the answer in the order of the reasoning steps; the second processing is used to make each reward value in the first reward value sequence meet a second set minimum difference requirement.

[0055] In this embodiment, the first processing is a first normalization process, used to eliminate excessive differences between the first reward value sequences corresponding to M answers, thereby ensuring the fairness and stability of subsequent relative advantage value calculation. The second processing is a second normalization process, used to eliminate excessive differences between the reward values ​​of N reasoning steps corresponding to the same answer, thereby ensuring the fairness and stability of subsequent relative advantage value calculation.

[0056] In this embodiment, the reward values ​​of each reasoning step in the answer are arranged in the order of the reasoning steps from front to back to form a second reward value sequence. Then, a second normalization process is performed on each reward value in the second reward value sequence to obtain a first reward value sequence.

[0057] Optionally, as an embodiment, the specific implementation of the second processing of the second reward value sequence corresponding to the answer to obtain the first reward value sequence corresponding to the answer can be as follows: obtain the mean and standard deviation of the reward values ​​of each reward value in the second reward value sequence of the answer. For each reasoning step of the answer, use the mean and standard deviation of the reward values ​​to perform a second normalization operation on the reward value corresponding to that reasoning step in the second reward value sequence of the answer, and obtain the second normalized reward value. The second normalized reward value of the reasoning step is the reward value corresponding to that reasoning step in the first reward value sequence. That is, the second normalized reward values ​​corresponding to each reasoning step are arranged in the order of the reasoning steps from front to back to form the first reward value sequence.

[0058] For example, the specific second normalization process is achieved through the following formula (3): Formula (3) Where ri_(norm,j) is the second normalized reward value for the j-th reasoning step of answering i; ri_j is the reward value for the j-th reasoning step of problem i; u is the average reward value of the N reasoning steps for problem i; σ is the standard deviation of the reward values ​​for the N reasoning steps of problem i.

[0059] Optionally, as an embodiment, the specific implementation of performing the first processing on the first reward value sequence corresponding to the M answers to obtain the processing result can be as follows: Obtain the global mean and standard deviation of the reward values ​​for each reward value in the first reward value sequence corresponding to M answers. For each reasoning step of each answer, use the global mean and standard deviation to perform a first normalization process on the reward value corresponding to that reasoning step in the first reward value sequence of that answer. The first normalized reward value of that reasoning step is the reward value corresponding to that reasoning step in the third reward value sequence. The first normalized reward values ​​corresponding to each reasoning step of the same answer are arranged in the order of the reasoning steps from front to back to form the third reward value sequence. The third reward value sequences of all answers constitute the above processing result.

[0060] Optionally, as an embodiment, the above processing result is a vector matrix, where any row or column in the vector matrix corresponds to a sequence of third reward values ​​for an answer.

[0061] S202, based on the processing results, as well as the weights and information entropy of each reasoning step in each answer, determine the relative advantage value of each answer.

[0062] Optionally, as an embodiment, for each answer, the weight and information entropy of each reasoning step in the answer, and the reward value corresponding to that reasoning step in the third reward value sequence corresponding to the answer in the processing result are subjected to a specified operation, and the relative advantage value of the answer is determined based on the operation result. For example, the weight and information entropy of the reasoning step and the reward value corresponding to that reasoning step in the third reward value sequence are multiplied, and the result of the multiplication operation is determined as the relative advantage value of the answer.

[0063] For example, the relative advantage value can be calculated using the following formula (4): Formula (4) in, The reasoning steps to answer i j The relative advantage value; The M relative advantage values ​​corresponding to answer i constitute the relative advantage value of answer i for question k, which is then processed by A. ki express; ri_(norm,j) The reasoning steps to answer i j The first normalized reward value; The reasoning steps to answer i j Linear weights; H j To answer the reasoning steps in i j Information entropy.

[0064] The above provides a detailed explanation of how the relative advantage value of this answer was determined.

[0065] To illustrate the method provided in this application in more detail, the following will be combined with... Figure 3 The solution provided in this application will be described in more detail by way of specific embodiments.

[0066] See Figure 3 , Figure 3 This is another flowchart illustrating the method provided in an embodiment of this application.

[0067] like Figure 3 As shown, the process may include the following steps: S301, for each of the S questions, the large language model outputs M answers based on the same input question. For each of the N reasoning steps of each answer, the reward value of that reasoning step is output using the PRM model.

[0068] S302, for each reasoning step, determine the linear weight of the reasoning step based on the step number of the reasoning step in the answer and the total number of reasoning steps in the answer.

[0069] Specifically, the above formula (1) can be used for calculation.

[0070] S303, for each reasoning step, perform a specified operation on the reward value of the reasoning step and the difference between the specified value and the reward value of the reasoning step, and determine the result of the specified operation as the information entropy of the reasoning step.

[0071] Specifically, the above formula (2) can be used for calculation.

[0072] S304, the reward values ​​of each reasoning step in the answer are arranged in the order of the reasoning steps from front to back to form a second reward value sequence. According to the above formula (3), the reward value of each reasoning step in the second reward value sequence of the answer is subjected to a second normalization process to obtain the second normalized reward value of the reasoning step. The second normalized reward values ​​corresponding to each reasoning step are arranged in the order of the reasoning steps from front to back to form a first reward value sequence.

[0073] Specifically, the second reward value sequence is a 1×N vector, and the resulting first reward value sequence is also a 1×N vector.

[0074] S305, based on the first reward value sequence corresponding to M answers, perform a first normalization process on the reward value corresponding to each reasoning step in the first reward value sequence of each answer to obtain a third reward value sequence.

[0075] Specifically, the sequence of third reward values ​​corresponding to the M answers forms an M*N matrix, with each row vector corresponding to the sequence of third reward values ​​for one answer.

[0076] S306. Based on the M*N matrix, the weights and information entropy of each reasoning step in the M answers, determine the relative advantage value of each answer.

[0077] Specifically, the relative advantage value of each answer is also a 1×N vector.

[0078] S307: Construct a loss function based on the relative advantage values ​​of the M answers to each of the S questions, update the policy parameters of the large language model according to the loss function, until the stopping optimization condition is met, and determine the large language model that meets the stopping optimization condition as the target large language model.

[0079] For example, the loss function can be constructed using the following formula (5): Formula (5) in, The loss function; s is the total number of problems in the current training batch; m is the total number of answers generated for each question; Let i be the relative advantage value corresponding to the answer i to question K; This is to determine the probability of the current large language model generating answer i for question k, and the change in the probability of the large language model generating answer i for question k compared to the previous batch of training, in order to prevent the large language model from suddenly changing under answer i and causing training to crash. γ: Represents the coefficient of the KL divergence penalty term, used to control the magnitude of policy parameter updates in large language models; γ : is the KL divergence, a conservative constraint to prevent the large language model from deviating excessively from the large language model before this batch of training in pursuit of high rewards.

[0080] The methods provided in the embodiments of this application have been described above. The apparatus provided in the embodiments of this application is described below: See Figure 4 , Figure 4 This is a structural diagram of the device provided in an embodiment of this application. Figure 4 As shown, the device is applied to a network access device and includes: a first determining module 401, a second determining module 402, a third determining module 403, and an optimization module 404.

[0081] The first determining module 401 is used to determine the weight of each reasoning step in each answer output by the large language model; the weight of any reasoning step in the answer is used to characterize the relative importance of the reasoning step to the correctness of the answer among all reasoning steps in the answer; the larger the weight of the reasoning step, the greater the relative importance of the reasoning step; the large language model outputs M answers based on the same input question; M is greater than 1; The second determining module 402 is used to determine the information entropy of each reasoning step using the reward value of each reasoning step in the answer; the information entropy of any reasoning step is used to represent the confidence level of the reward value of the reasoning step; the smaller the information entropy of any reasoning step, the greater the confidence level of the reward value of the reasoning step. The third determining module 403 is used to determine the relative advantage value of the answer based on the reward value, weight, and information entropy of each reasoning step in each answer. The relative advantage value of any answer is used to represent the relative value of the answer among M answers. Optimization module 404 is used to optimize the large language model based on the relative advantage value of each answer.

[0082] As an example, when the first determining module 401 performs the step of determining the weight of each reasoning step in the answer, it is specifically used for: For each reasoning step in the answer, the linear weight of the reasoning step is determined based on the step number of the reasoning step in the answer and the total number of reasoning steps in the answer.

[0083] As an example, when the second determining module 402 performs the step of determining the information entropy of the reasoning step using the reward value of each reasoning step in the answer, it is specifically used for: For each inference step, the information entropy of that inference step is determined based on the reward value of that inference step and the difference between the specified value and the reward value of that inference step.

[0084] As an example, when the third determining module 403 performs the step of determining the relative advantage value of the answer based on the reward value, weight, and information entropy of each reasoning step in each answer, it is specifically used for: A first processing step is performed on the first reward value sequence corresponding to M answers to obtain a processing result. The first processing step is used to ensure that each reward value in the processing result meets a first set minimum difference requirement. The first reward value sequence corresponding to any answer is obtained by performing a second processing step on the second reward value sequence corresponding to that answer. The second reward value sequence corresponding to any answer is a sequence formed by the reward values ​​of each reasoning step in that answer according to the order of the reasoning steps. The second processing step is used to ensure that each reward value in the first reward value sequence meets a second set minimum difference requirement. Based on the processing results, as well as the weights and information entropy of each reasoning step in each answer, the relative advantage value of each answer is determined.

[0085] As an example, when the third determining module 403 performs the second processing on the second reward value sequence corresponding to any answer to obtain the first reward value sequence corresponding to that answer, it is specifically used for: Obtain the mean and standard deviation of the reward values ​​for each reward value in the second reward value sequence for this answer; For each reasoning step of the answer, the reward value corresponding to the reasoning step in the first reward value sequence is obtained based on the reward value, the mean reward value, and the standard deviation of the reward value corresponding to the reasoning step in the second reward value sequence of the answer.

[0086] As an example, when the third determining module 403 performs the first processing on the first reward value sequence corresponding to the M answers to obtain the processing result, it is specifically used for: Obtain the global reward mean and global reward standard deviation of each reward value in the first reward value sequence corresponding to M answers; For each reasoning step of each answer, the reward value corresponding to that reasoning step in the processing result is obtained based on the reward value corresponding to that reasoning step in the first reward value sequence of the answer, the mean of global reward values, and the standard deviation of global reward values.

[0087] As an example, the processing result is a vector matrix, where any row or column corresponds to a third reward value sequence for an answer. This third reward value sequence is obtained by processing the first reward value sequence for the answer. The third determining module 403, when performing the step of determining the relative advantage value of each answer based on the processing results and the weights and information entropies of each reasoning step in each answer, is specifically used for: For each answer, the weight and information entropy of each reasoning step in the answer, and the reward value corresponding to that reasoning step in the third reward value sequence of the answer in the processing result are calculated accordingly; the relative advantage value of the answer is determined based on the calculation result.

[0088] This concludes the process. Figure 4 Structural description of the device shown.

[0089] See Figure 5 , Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application. Figure 5 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.

[0090] Based on the same concept as the above-described method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above-described examples of this application.

[0091] For example, the aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0092] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A question-answering method based on a large language model, characterized in that, The method includes: For each answer output by the large language model, the weight of each reasoning step in that answer is determined. The weight of any reasoning step in the answer represents the relative influence of that reasoning step on the quality of the answer among all reasoning steps in that answer. The larger the weight of the reasoning step, the greater its relative influence on the quality of the answer. The large language model outputs M answers based on the same input question, where M is greater than 1. The information entropy of each reasoning step is determined by the reward value of that reasoning step in the answer. The information entropy of any reasoning step is related to the confidence of the reward value of that reasoning step. The larger the information entropy of any reasoning step, the smaller the confidence of the reward value of that reasoning step. Based on the reward value, weight, and information entropy of each reasoning step in each answer, the relative advantage value of the answer is determined. The relative advantage value of any answer is used to represent the relative value of the answer among M answers. The large language model is optimized based on the relative advantage value of each answer.

2. The method according to claim 1, characterized in that, The determination of the weight of each reasoning step in the answer includes: For each reasoning step in the answer, the linear weight of the reasoning step is determined based on the step number of the reasoning step in the answer and the total number of reasoning steps in the answer.

3. The method according to claim 1, characterized in that, The process of determining the information entropy of each reasoning step using the reward value of that answer includes: For each inference step, the information entropy of that inference step is determined based on the reward value of that inference step and the difference between the specified value and the reward value of that inference step.

4. The method according to claim 1, characterized in that, The determination of the relative advantage value of an answer based on the reward value, weight, and information entropy of each reasoning step in each answer includes: A first processing step is performed on the first reward value sequence corresponding to M answers to obtain a processing result; the first processing step is used to ensure that each reward value in the processing result meets a first set minimum difference requirement; wherein, the first reward value sequence corresponding to any answer is obtained by performing a second processing step on the second reward value sequence corresponding to that answer; the second reward value sequence corresponding to any answer is a sequence formed by the reward values ​​of each reasoning step in that answer in the order of the reasoning steps; the second processing step is used to ensure that each reward value in the first reward value sequence meets a second set minimum difference requirement; Based on the processing results, as well as the weights and information entropy of each reasoning step in each answer, the relative advantage value of each answer is determined.

5. The method according to claim 4, characterized in that, The second processing of the second reward value sequence corresponding to any answer yields the first reward value sequence corresponding to that answer, including: Obtain the mean and standard deviation of the reward values ​​for each reward value in the second reward value sequence for this answer; For each reasoning step in the answer, the reward value corresponding to that reasoning step in the first reward value sequence is obtained based on the reward value corresponding to that reasoning step in the second reward value sequence of the answer, the mean of the reward values, and the standard deviation of the reward values.

6. The method according to claim 4, characterized in that, The first processing of the sequence of first reward values ​​corresponding to the M answers yields the following results: Obtain the global reward mean and global reward standard deviation of each reward value in the first reward value sequence corresponding to M answers; For each reasoning step of each answer, the reward value corresponding to that reasoning step of the answer in the processing result is obtained based on the reward value corresponding to that reasoning step in the first reward value sequence of the answer, the mean of the global reward value, and the standard deviation of the global reward value.

7. The method according to claim 4, characterized in that, The processing result is a vector matrix, where any row or column of the vector matrix corresponds to a third reward value sequence for a given answer. This third reward value sequence is obtained by processing the first reward value sequence for the given answer. The determination of the relative advantage value of each answer based on the processing result, as well as the weight and information entropy of each reasoning step in each answer, includes: For each answer, the weight and information entropy of each reasoning step in the answer, and the reward value corresponding to that reasoning step in the third reward value sequence corresponding to the answer in the processing result are calculated; the relative advantage value of the answer is determined based on the calculation result.

8. A question-answering device based on a large language model, characterized in that, The device includes: The first determining module is used to determine the weight of each reasoning step in each answer output by the large language model. The weight of any reasoning step in the answer is used to characterize the relative importance of the reasoning step to the correctness of the answer among all reasoning steps in the answer. The larger the weight of the reasoning step, the greater the relative importance of the reasoning step. The large language model outputs M answers based on the same input question, where M is greater than 1. The second determining module is used to determine the information entropy of each reasoning step in the answer by using the reward value of each reasoning step; the information entropy of any reasoning step is used to represent the confidence level of the reward value of the reasoning step; the smaller the information entropy of any reasoning step, the greater the confidence level of the reward value of the reasoning step. The third determining module is used to determine the relative advantage value of each answer based on the reward value, weight, and information entropy of each reasoning step in each answer. The relative advantage value of any answer is used to represent the relative value of the answer among M answers. An optimization module is used to optimize the large language model based on the relative advantage value of each answer.

9. The apparatus according to claim 8, characterized in that, When the first determining module performs the step of determining the weight of each reasoning step in the answer, it is specifically used for: For each reasoning step in the answer, the linear weight of the reasoning step is determined based on the step number of the reasoning step in the answer and the total number of reasoning steps in the answer. And / or, When the second determining module performs the step of determining the information entropy of each reasoning step using the reward value of the answer, it is specifically used for: For each inference step, the information entropy of that inference step is determined based on the reward value of that inference step and the difference between the specified value and the reward value of that inference step. And / or, When the third determining module performs the step of determining the relative advantage value of an answer based on the reward value, weight, and information entropy of each reasoning step in each answer, it is specifically used for: A first processing step is performed on the first reward value sequence corresponding to M answers to obtain a processing result; the first processing step is used to ensure that each reward value in the processing result meets a first set minimum difference requirement; wherein, the first reward value sequence corresponding to any answer is obtained by performing a second processing step on the second reward value sequence corresponding to that answer; the second reward value sequence corresponding to any answer is a sequence formed by the reward values ​​of each reasoning step in that answer in the order of the reasoning steps; the second processing step is used to ensure that each reward value in the first reward value sequence meets a second set minimum difference requirement; Based on the processing results, as well as the weights and information entropy of each reasoning step in each answer, the relative advantage value of each answer is determined; And / or, When the third determining module performs the second processing on the second reward value sequence corresponding to any answer to obtain the first reward value sequence corresponding to that answer, it is specifically used for: Obtain the mean and standard deviation of the reward values ​​for each reward value in the second reward value sequence for this answer; For each reasoning step in the answer, the reward value corresponding to the reasoning step in the first reward value sequence is obtained based on the reward value corresponding to the reasoning step in the second reward value sequence of the answer, the mean of the reward value, and the standard deviation of the reward value. And / or, When the third determining module performs the step of performing the first processing on the first reward value sequence corresponding to the M answers to obtain the processing result, it is specifically used for: Obtain the global reward mean and global reward standard deviation of each reward value in the first reward value sequence corresponding to M answers; For each reasoning step of each answer, the reward value corresponding to that reasoning step of the answer in the processing result is obtained based on the reward value corresponding to that reasoning step in the first reward value sequence of the answer, the mean of the global reward value, and the standard deviation of the global reward value. And / or, The processing result is a vector matrix, where any row or column of the vector matrix corresponds to a third reward value sequence for a given answer. This third reward value sequence is obtained by processing the first reward value sequence for the given answer. When the third determining module performs the step of determining the relative advantage value of each answer based on the processing result and the weights and information entropy of each reasoning step in each answer, it is specifically used for: For each answer, the weight and information entropy of each reasoning step in the answer, and the reward value corresponding to that reasoning step in the third reward value sequence corresponding to the answer in the processing result are calculated; the relative advantage value of the answer is determined based on the calculation result.

10. An electronic device, characterized in that, The electronic device includes: Processor; and A machine-readable storage medium storing machine-executable instructions that, when executed by the processor, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.

11. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions that, when executed by a processor, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent agent training method and device, equipment and storage medium

    CN119494383A

  • Optimization method and device for improving near-end strategy based on language model, and electronic equipment

    CN120068993A

  • Model training method and device, task execution method and device, electronic equipment and storage medium

    CN120930803A

  • Large reasoning model factuality enhancement method, system and equipment based on reinforcement learning and medium

    CN121457634A

  • Service execution method and device based on information verification, medium and equipment

    CN121638369A