A method, apparatus, storage medium, and electronic device for generating response information.
By comparing the current response information with historical response information of a large language model, and combining iteration and resampling, the response information generation process is optimized, which solves the accuracy and robustness problems of the large language model and improves the accuracy and stability of the response information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
- Filing Date
- 2024-11-22
- Publication Date
- 2026-07-31
AI Technical Summary
Existing large language models have low accuracy and poor robustness in generating response information, and suffer from overconfidence.
By acquiring the user's input question to be answered, the current response information is generated using a large language model and compared with several historical response information to determine the response information to be compared. The difference result of the current response is calculated, and the difference result is input into the model for iteration until the stopping condition is met. Combined with resampling and confidence calibration, the response information is optimized.
This improved the accuracy and stability of the output response information of the large language model, reduced overconfidence, and enhanced the robustness of the model.
Smart Images

Figure CN119202204B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computers, and in particular to a method, apparatus, storage medium, and electronic device for generating response information. Background Technology
[0002] The large language model obtained through artificial intelligence technology can determine the response information based on various questions raised by users and display it to them. Among these questions, private information may be included.
[0003] Current large language models have low accuracy in determining response information, suffer from overconfidence, and exhibit poor robustness.
[0004] Based on this, this specification provides a method for generating response information. Summary of the Invention
[0005] This specification provides a method, apparatus, storage medium, and electronic device for generating response information, to at least partially solve the aforementioned problems existing in the prior art.
[0006] The following technical solution is adopted in this specification:
[0007] This specification provides a method for generating response information, including:
[0008] Get the user's input for the question to be answered;
[0009] The question to be answered is input into a large language model to obtain the current response information output by the large language model; and several historical response information output by the large language model based on the question to be answered is obtained.
[0010] Identify the response information to be compared from a number of historical response messages;
[0011] Based on the current response information and the response information to be compared, determine the difference result of the current response; and determine the current response information as historical response information;
[0012] The current response difference result and the question to be responded to are input into the large language model to redetermine the current response information until the preset stopping iteration condition is met.
[0013] Retrieve and display the current response information when iteration stops.
[0014] Optionally, the question to be answered is input into a large language model to obtain the current response information output by the large language model, specifically including:
[0015] Get the preset first prompt message;
[0016] Based on the preset first prompt information and the question to be answered, resampling is performed to obtain the first resampling result;
[0017] The first resampling result is input into the large language model to obtain the current response information output by the large language model.
[0018] Optionally, before determining the difference result of the current response based on the current response information and the response information to be compared, the method further includes:
[0019] The current response information is resampled to obtain the resampled current response information.
[0020] Optionally, the difference result of the current response is determined based on the current response information and the response information to be compared, specifically including:
[0021] Obtain several historical response difference results to obtain a historical difference set;
[0022] Based on the aforementioned set of historical differences, the second prompt message is determined;
[0023] Based on the second prompt information, the current response information is compared with the response information to be compared, and the difference result of the current response is determined.
[0024] Optionally, based on the second prompt information, the current response information is compared with the response information to be compared to determine the difference result of the current response, specifically including:
[0025] Based on the second prompt information, the current response information is compared with the response information to be compared, and the feedback result is determined;
[0026] Based on the feedback results and the historical difference set, the current response difference result is determined.
[0027] Optionally, the current response difference result and the question to be responded to are input into the large language model to redetermine the current response information, specifically including:
[0028] Based on the current response discrepancy results, determine the third prompt message;
[0029] Based on the third prompt information and the question to be answered, resampling is performed to obtain the second resampling result;
[0030] The second resampling result and the current response difference result are input into the large language model to redetermine the current response information.
[0031] Optionally, the current response information can be redefined, specifically including:
[0032] When the current response information is no different from the response to be compared, resampling is performed based on the preset first prompt information and the question to be answered to obtain the first resampling result;
[0033] The first resampling result is input into the large language model to redetermine the current response information.
[0034] Optionally, the method further includes:
[0035] All response information generated during the iteration process is obtained to obtain a response information set, which includes historical response information, current response information, and response information to be compared.
[0036] Determine the response distribution result of the response information set;
[0037] Based on the response distribution results, determine the confidence level of the current response information when stopping iteration, and display it.
[0038] This specification provides a response information generation device, the device comprising:
[0039] The acquisition module is used to acquire the user's input questions that need to be answered;
[0040] The response information determination module is used to input the question to be answered into the large language model, obtain the current response information output by the large language model, and acquire several historical response information output by the large language model based on the question to be answered.
[0041] The comparison information determination module is used to determine the reply information to be compared from a number of historical reply information;
[0042] The difference determination module is used to determine the difference result of the current response based on the current response information and the response information to be compared; and to determine the current response information as historical response information.
[0043] The iteration module is used to input the current response difference result and the question to be responded to into the large language model, redetermine the current response information, until the preset stop iteration condition is met;
[0044] The display module is used to obtain and display the current response information when iteration stops.
[0045] Optionally, the response information determination module is specifically used to obtain preset first prompt information; based on the preset first prompt information and the question to be responded to, perform resampling to obtain a first resampling result; input the first resampling result into a large language model to obtain the current response information output by the large language model.
[0046] Optionally, the device further includes:
[0047] The resampling module is used to resample the current response information to obtain the resampled current response information.
[0048] Optionally, the difference determination module is specifically used to obtain several historical response difference results to obtain a historical difference set; determine a second prompt message based on the historical difference set; and compare the current response information with the response information to be compared according to the second prompt message to determine the current response difference result.
[0049] Optionally, the difference determination module is specifically used to compare the current response information with the response information to be compared according to the second prompt information, and determine the feedback result; and determine the current response difference result according to the feedback result and the historical difference set.
[0050] Optionally, the iterative module is specifically used to determine a third prompt based on the current response difference result; to resample based on the third prompt and the question to be answered to obtain a second resampling result; and to input the second resampling result and the current response difference result into the large language model to redetermine the current response information.
[0051] Optionally, the iteration module is specifically used to, when there is no difference between the current response information and the response to be compared, perform resampling based on the preset first prompt information and the question to be answered, to obtain a first resampling result; input the first resampling result into the large language model to redetermine the current response information.
[0052] Optionally, the device further includes:
[0053] The confidence determination module is used to acquire all response information generated during the iteration process to obtain a response information set, which includes historical response information, current response information, and response information to be compared; determine the response distribution result of the response information set; and determine and display the confidence level of the current response information when the iteration stops based on the response distribution result.
[0054] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described response information generation method.
[0055] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described response information generation method.
[0056] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0057] In the response information generation method provided in this specification, a user-inputted question is obtained, and the question is input into a large language model to obtain the current response information output by the large language model. Several historical response messages output by the large language model based on the question are also obtained, and a response message to be compared is determined from these historical responses. Based on the current response information and the response message to be compared, a difference result is determined for the current response, and the current response information is designated as historical response information. The current response difference result and the question are input into the large language model to re-determine the current response information until a preset stopping iteration condition is met. The current response information at the point of stopping iteration is then obtained and displayed.
[0058] As can be seen from the above method, by comparing historical and current responses, the differences in responses are identified. These differences, along with the question to be answered, are then input back into the large language model. This prompts the model to pay attention to these differences, reducing overconfidence and improving the accuracy of the responses output by the large language model. Furthermore, since the responses to be compared are historical, using them as input for the next iteration also improves the stability of the responses output by the large language model, preventing excessive differences in responses between iterations and enhancing the robustness of the large language model. Attached Figure Description
[0059] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0060] Figure 1 This is a flowchart illustrating a method for generating response information provided in this specification.
[0061] Figure 2 This is a schematic diagram of the self-consistent decoding strategy provided in this specification;
[0062] Figure 3 This is a flowchart illustrating another method for generating response information provided in this specification.
[0063] Figure 4 This is a schematic diagram of the second prompt information provided in this instruction manual;
[0064] Figure 5 This is a diagram illustrating the third prompt information provided in this instruction manual;
[0065] Figure 6 This is a schematic diagram of a response information generation device provided in this specification;
[0066] Figure 7The corresponding information provided in this specification Figure 1 A schematic diagram of an electronic device. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this application.
[0068] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0069] Figure 1 This is a flowchart illustrating a method for generating response information provided in this specification, which specifically includes the following steps:
[0070] S100: Obtain the user-inputted question awaiting a response.
[0071] Large language models can generate response information based on user-input questions, leveraging their powerful reasoning capabilities. The number of parameters in a large language model is at least in the hundreds of millions. For example, in a knowledge-based question-answering scenario, the user inputs several questions, and the large language model outputs the answers. Similarly, in a mathematical logic reasoning scenario, the user inputs a mathematical formula to be reasoned, and the large language model outputs the reasoning result. However, current large language models still have some problems, including low accuracy of the output response information, overconfidence, and poor robustness. Understandably, overconfidence not only provides incorrect response information but also has a high confidence level in incorrect predictions. Therefore, this specification provides a method for generating response information. The execution entity of this specification can be a computing device that deploys a large language model to generate response information, such as a server, or a device such as a desktop computer or laptop computer; this specification does not impose any restrictions. For ease of description, the following explanation focuses on a server as the execution entity for the response information generation method provided in this specification.
[0072] The server needs to obtain the user's pending questions in order to determine the appropriate response. Specifically, in response to user input, it obtains the pending questions entered by the user.
[0073] S102: Input the question to be answered into the large language model to obtain the current response information output by the large language model; and obtain several historical response information output by the large language model based on the question to be answered.
[0074] In one or more embodiments of this specification, after obtaining the information to be replied to, the server can input the information to be replied to into the large language model and obtain the current reply information output by the large language model.
[0075] Generally, the large language model is a pre-trained language model with response capabilities. Specifically, whether a large language model is well-trained can be defined based on criteria such as the accuracy of its output response information. For example, if the accuracy of the large language model is 65% or higher, it is defined as well-trained; otherwise, it is considered untrained. For large language models with low accuracy, they can be trained first to improve accuracy, and then the response information generation method provided in this manual can be used to further improve the accuracy of the response information. Alternatively, the response information generation method provided in this manual can be used directly, without involving the training process of the large language model during response information generation.
[0076] Furthermore, the server can also obtain several historical response messages output by the large language model based on the information to be responded to, which are used to subsequently determine the prompt message for the large language model. Understandably, during each iteration, the server inputs the information to be responded to into the large language model, obtains the current response message, and before proceeding to the next iteration, identifies the current response message as a historical response message. For example, in the nth iteration, the current response message is obtained as 1; subsequently, before the (n+1)th iteration, this current response message n is identified as the historical response message n-1.
[0077] It should be noted that in the nth iteration, the historical response information can include all response information generated during the first to the (n-1)th iteration. Therefore, when retrieving historical response information, the server can retrieve only a preset number of historical responses, which can be set as needed; this specification does not impose any restrictions on this. The server can also determine a number of historical responses based on preset conditions. These preset conditions may include a similarity between the determined historical responses exceeding a preset threshold. Specifically, feature analysis can be performed on all historical responses, and for each historical response, its similarity to the remaining historical responses excluding itself can be determined.
[0078] S104: Identify the response information to be compared from a number of historical response messages.
[0079] Specifically, the server can use a voting mechanism to select responses with high consistency from several historical responses as the responses to be compared. Consistency refers to the consistency of the content in the responses. For example, if a user enters a piece of text and is asked to determine the types of animals appearing in that text, and three historical responses are obtained, all three responses list three types of animals. However, the first response lists dogs, cats, and rabbits; the second lists dogs, sheep, and pigs; and the third lists dogs, sheep, and cows. Through the voting mechanism, the responses to be compared can be determined to list three types of animals: dogs, sheep, and pigs.
[0080] Of course, other methods can also be used to determine the response information to be compared from several historical response information, including feature comparison methods, to determine the response information with high feature consistency among the historical response information. Specifically, the feature analysis results of the determined historical response information in step S104 can be obtained, and the final response information can be determined based on the feature analysis results.
[0081] Using historical responses with high consistency as the responses to be compared ensures the stability of the output of the large language model. The large language model will not produce responses with huge differences due to small changes in subsequent input information, which improves the robustness of the large language model.
[0082] S106: Determine the difference result of the current response based on the current response information and the response information to be compared; and determine the current response information as historical response information.
[0083] Figure 2 This is a schematic diagram of the self-consistent decoding strategy provided in this specification, such as... Figure 2 As shown.
[0084] In traditional self-consistent decoding strategies, the server inputs the question to be answered into a large language model, obtains n responses, and uses a voting mechanism to determine the final response from these n responses, which is then displayed. During the voting process, the large language model focuses on responses with high consistency, adhering to the logic of majority rule. This neglects opinions that differ from the majority, causing the large language model to completely ignore valuable information from minority opinions, leading to overconfidence.
[0085] This specification also uses a voting mechanism. Therefore, to address the issue of overconfidence while ensuring the robustness of the large language model, this specification can also determine the difference result of the current response based on the current response information and the response information to be compared. That is, the current response information is compared with the response information to be compared to determine the difference between the current response information and the response information to be compared, thus obtaining the current difference result. This current difference result is then used as subsequent input, causing the large language model to focus on the difference information, that is, to value the valuable information in the minority opinions, thereby eliminating the problem of overconfidence.
[0086] In addition, in order to continuously obtain the current reply information for comparison during subsequent iterations, the server also needs to identify the current reply information as historical reply information.
[0087] S108: Input the current response difference result and the question to be responded to into the large language model, redetermine the current response information, until the preset stopping iteration condition is met.
[0088] In one or more embodiments of this specification, the server may input the current response difference result and the question to be responded to into the large language model, obtain the output result of the large language model, and re-determine it as the current response information until a preset stopping iteration condition is met. The preset stopping iteration condition may be that the iteration stops when the number of iterations reaches a preset number, or it may be that the iteration stops when there is no difference between the current response information and the response information to be compared, etc. This specification does not limit this.
[0089] Traditional self-reflection methods struggle to identify errors in the large language model's own responses. This is because traditional methods rely solely on feedback from individual responses, making it difficult to pinpoint hidden errors and potentially generating misleading suggestions that ultimately result in inaccurate outputs. This specification addresses this by iteratively acquiring historical responses and identifying those to compare against the current response. This allows the large language model to reflect on multiple responses, particularly the differences between the current and historical responses. This enables the model to identify potential errors in its output, thereby improving accuracy.
[0090] S110: Obtain and display the current response information when iteration stops.
[0091] For example, if the iteration stops when the number of iterations reaches 4, then the response output of the large language model includes historical response information 1~3 and current response information 4. The server can obtain the current response information 4 and display it to the user.
[0092] based on Figure 1The method for generating response information shown compares historical responses with current responses to identify differences. These differences, along with the question to be answered, are then input back into the large language model. This prompts the model to pay attention to these differences, reducing overconfidence and improving the accuracy of the responses output. Furthermore, since the responses to be compared are historical, using them as input for the next iteration also improves the stability of the responses output by the large language model, preventing excessive differences in responses between iterations and enhancing the model's robustness.
[0093] Figure 3 This is a flowchart illustrating another method for generating response information provided in this specification, such as... Figure 3 As shown.
[0094] Regarding step S102, when inputting the question to be answered into the large language model, the server can also obtain a preset first prompt message, which is determined based on the application scenario of the large language model. For example, if the application scenario is a mathematical logic reasoning scenario, the constructed first prompt message needs to include mathematical scenario knowledge to prompt the large language model. For example, the first prompt message could be "Please deduce the result obtained from formula x step by step." The presence of the prompt message can further improve the accuracy of the response information output by the large language model.
[0095] Afterwards, the server can resample based on the preset first prompt information and the question to be answered, to obtain the first resampled result. Resampling is mainly for the purpose of subsequently correcting and optimizing the generated response information, ensuring that the sequence of generated response information is consistent and coherent in terms of syntax, logic, and semantics, and also helps to further solve the illusion problem existing in traditional self-reflection methods.
[0096] Then, the server can input the first resampling result into the large language model to obtain the current response information output by the large language model. Figure 3In step S104, based on the first prompt information, the question to be answered is resampled to obtain the first resampling result, which is then input into the large language model to obtain the current response information 1. If the current response information 1 is the response information obtained in the first iteration, then since there is no historical response information in the first iteration, there is no need to determine the response information to be compared and compare it with the current response information 1 to obtain the current response difference result. Step S104 is executed again, and the second iteration obtains the current response information 2. The historical response information is the current response information 1. Since the historical response information in the second iteration only includes the current response information 1, the response information to be compared obtained by voting is still the current response information 1. Then, in step S106, the response information to be compared 1 is compared with the current response information 2 to obtain the current response difference result 1. Similarly, in the third iteration, the response information to be compared 2 is compared with the current response information 3 to obtain the current response difference result 2. In the fourth iteration, the response information to be compared 3 is compared with the current response information 4 to obtain the current response difference result 3.
[0097] Before executing step S106, the server can resample the current response information to obtain resampled current response information. Then, when executing step S106, the server compares the resampled current response information with the response information to be compared to determine the difference in the current response.
[0098] In step S106, the server can also obtain prompting information, highlighting the key points of comparison for the large language model during the comparison process, in order to further improve the accuracy of the response information.
[0099] Specifically, the server can obtain several historical response difference results to form a historical difference set. Based on this historical difference set, a second prompt message is determined. Specifically, this can be a feature whose proportion in the historical difference set reaches a preset percentage. The second prompt message is determined based on this feature. For example, if a certain feature consistently exists in the historical difference set, indicating a problem with the large language model's output of that feature, the second prompt message could be that feature, causing the large language model to focus on whether there is still a difference between the current response and the response to be compared regarding that feature when determining response differences. If this feature is the number of animal species, meaning the server determines that each historical response difference result includes the number of animal species, it means the number of animal species determined by the large language model is always different. Therefore, the second prompt message could be "Please pay attention to the number of animal species in the question to be answered."
[0100] Then, the server can compare the current response with the response to be compared based on the second prompt information to determine the current response difference result. This specification does not limit the number of historical response difference results in the historical difference set, and it is understood that the historical response difference results are the response difference results determined in any iteration process before the current iteration.
[0101] Figure 3 In the comparison between the current response information 3 and the response information 2 to be compared, the current response difference result 1 pointing to the current response difference result 2 indicates that the server can determine the prompt information 2 based on the current response difference result 1, and then compare it with the current response difference result 2 to obtain the current response difference result 2. Similarly, the prompt information 3 can be determined based on at least one of the current response difference result 1 and the current response difference result 2. Prompt information 1 to 3 are the second prompt information in this specification. Since there is no response difference result in the first iteration, in the second iteration, prompt information 1 can be a preset prompt template. This specification does not specifically limit the specific content of the preset prompt template, and it can be derived from experience.
[0102] Based on prompt message 1, the server compares the reply message to be compared 1 with the current reply message 2, obtaining the current reply difference result 1. Similarly, based on prompt message 2, the server compares the reply message to be compared 2 with the current reply message 3, obtaining the current reply difference result 2. Based on prompt message 3, the server compares the reply message to be compared 3 with the current reply message 4, obtaining the current reply difference result 3.
[0103] Furthermore, the server can also compare the current response with the response to be compared based on the second prompt information to determine the feedback result. Based on the feedback result and the historical difference set, the server determines the current response difference result. In other words, the difference between the current response and the response to be compared is used as the feedback result. This feedback result is then fused with several historical response difference results from the historical difference set to obtain a difference fusion result, which is then determined as the current response difference result.
[0104] The fusion operation allows the determined difference in the current response to include the differences in previous iterations. Therefore, in the next iteration, the differences from any previous iteration are still retained. This allows the large language model to check whether it still has the problems that existed in previous iterations based on the response differences, further improving the accuracy of the output response information and solving the problem of overconfidence.
[0105] Figure 4 This is a schematic diagram of the second prompt information provided in this instruction manual, such as... Figure 4 As shown.
[0106] The second prompt message is as follows: "Given two alternative solutions to a problem, carefully analyze and compare the differences in their reasoning steps, and reflect on: 1) the specific differences in their reasoning steps and final response information; 2) the reasons for these differences."
[0107] Pending questions: [Pending questions]
[0108] Two solutions:
[0109] Solution 1: [Majority voting]
[0110] Solution 2: [Current Reply Information]
[0111] If no differences exist, stop the iteration. If differences exist, describe the differences, identify the errors in the response information, and explain the reasons. Determine key recommendations to prevent such errors and combine them with the response difference results to obtain new response difference results.
[0112] Figure 4 This example uses the condition of stopping iteration when there is no difference between the current response and the response to be compared. Figure 4 Solution 1, majority voting, actually refers to the information to be compared and responded to after being determined through a voting medium. Feedback refers to executing step S108, which involves inputting the current response difference result and the question to be responded to into the large language model to redetermine the current response information. It should be noted that when using the second prompt information, it is necessary to... Figure 4 The square brackets should be replaced with specific content relevant to the scenario. For example, if the question to be answered is "Please determine the value of x in the following formula based on the formula and specific values," then the question to be answered in the second prompt message is the content of the aforementioned question to be answered.
[0113] Therefore, during step S108, the server can also determine the third prompt information based on the current response difference result, including extracting features from the current response difference result, or directly using the current response difference result as the third prompt information. Then, based on the third prompt information and the question to be answered, resampling is performed to obtain the second resampling result. Finally, the server can input the second resampling result and the current response difference result into the large language model to redetermine the current response information. Figure 3 In the middle, the current response difference result 1 and the current response difference result 2 are used as the third prompt information.
[0114] Figure 5 This is a diagram illustrating the third prompt information provided in this instruction manual, such as... Figure 5 As shown.
[0115] The third prompt message reads: "Solve the following problems step by step, starting each step with 'Step' and ensuring that each step is separated by '\n\n' and ending with 'So the reply message is', followed by the reply message."
[0116] Consider integrating the previous response discrepancies into your solution process.
[0117] Pending questions: [Pending questions]
[0118] Reply message
[0119] Furthermore, if the preset stopping iteration condition is not that there is no difference between the current response information and the response information to be compared, when the iteration stops, if there is no difference between the current response information and the response to be compared, resampling is performed based on the preset first prompt information and the question to be answered to obtain the first resampling result. The first resampling result is then input into the large language model to redetermine the current response information.
[0120] It should be further noted that this specification also allows for the calibration of the confidence level of the large language model's output. Confidence calibration of the large language model refers to a reasonable measurement of the uncertainty of its own output. An ideal large language model should have a reasonable estimate of the accuracy of its own answers, thereby alerting users to unreliable answers in low-confidence scenarios and improving the output with the help of external resources. Understandably, in this specification, the large language model can utilize various prompts for improvement. This method of using prompts and iteratively improving the output can also be called chain reasoning. By dividing complex tasks into a series of logical reasoning steps to simulate the human reasoning process, it stimulates the potential reasoning ability of the large language model. This method provides a structured mechanism for using large language models for complex reasoning tasks.
[0121] When calibrating the confidence level, the server can obtain all response information generated during the iteration process, resulting in a response information set. This set includes historical response information, current response information, and response information to be compared. Generally, the server can perform confidence level calibration after obtaining the current response at the time the iteration stops. The response information set includes several historical response information, one current response information, and one response information to be compared. Alternatively, confidence level calibration can be performed once during each iteration; this specification does not restrict this approach.
[0122] The server can then determine the response distribution of the response set. This distribution can be the proportion of each feature in the responses or the number of identical responses. For example, if there are three existing responses with 3, 4, and 3 animal species respectively, the response distribution can be determined to be 66.6%. Since the current response information in this specification is determined based on the differences in historical response information, the accuracy of the response information is improved. Consequently, the confidence level determined based on the response distribution is also more accurate, avoiding erroneous predictions and providing higher confidence levels.
[0123] Finally, the server can determine and display the confidence level of the current response information when iteration stops based on the response distribution result. The response distribution result can be directly used as the confidence level of the current response information when iteration stops, or statistical analysis can be performed on the response distribution result to obtain the analysis results, which can then be used as the confidence level of the current response information when iteration stops. This specification does not impose any restrictions on this approach.
[0124] Traditional sampling-based confidence calibration methods rely on "independent" repeated sampling, neglecting the analysis of differences between responses. In contrast, this specification leverages the self-reflective ability of large language models to address response discrepancies to improve their confidence calibration capabilities. Answers that remain consistent after multiple comparisons and reflections are more reliable, while responses that change frequently after comparisons and reflections reflect the inherent uncertainty of large language models. Therefore, the distribution of multiple large language models generated based on iterative reflection is used as a measure of model confidence.
[0125] In summary, this specification combines a self-consistent decoding strategy with the self-reflection capability of a large language model, thereby improving the accuracy of reasoning decisions in the large language model and mitigating overconfidence in the confidence assessment process, thus overcoming the limitations of traditional self-consistent decoding strategies and self-reflection methods.
[0126] To address the issue that traditional self-consistent decoding methods for large language models completely ignore valuable information from minority opinions, this specification iteratively compares the differences between minority and majority opinions to identify inconsistencies in the resampling process and the potential errors and illusions they contain. This, in turn, prompts large language models to be more mindful of avoiding potential errors in subsequent generation processes and mitigates overconfidence in confidence estimation through reflection on inconsistencies.
[0127] To address the issue of insufficient robustness in traditional self-reflection methods, this specification enhances the robustness of reflection results by introducing a self-consistent logic of multiple iterations and a final vote. Furthermore, to address the difficulty of traditional self-reflection methods in identifying errors, this specification finds potential errors by generating multiple response messages and comparing them. This method effectively addresses the issues of overconfidence and insufficient robustness of large language models when searching for potential errors.
[0128] Through comprehensive experimental comparisons, the response information generation method (referred to as Mirror-Consistency) provided in this manual demonstrates significantly superior performance compared to the traditional self-consistency approach. The experiments used four large language models of different sizes and from different sources: gpt-3.5-turbo, qwen-turbo, llama3-8B, and llama3-70B, to ensure the generalizability of the experiments. Four open-source datasets were used for evaluation: GSM8K, SVAMP, StrategyQA, and Date. The first two are mathematical logic reasoning tasks, while the latter two are multi-hop knowledge-based question-answering tasks. The experimental results show that the response information generation method provided in this manual improves reasoning accuracy in the vast majority of tasks, especially those involving complex multi-step mathematical reasoning.
[0129] In addition, the performance of the response information generation method provided in this specification and the self-consistency scheme was compared on the confidence assessment task. The experimental results show that the calibration curve corresponding to the response information generation method provided in this specification is closer to the ideal curve in most cases. Especially when the traditional self-consistency scheme exhibits overconfidence problems, such as when using qwen-turbo as the evaluation model, it further demonstrates that the response information provided in this specification can significantly improve the overconfidence problem of large language models.
[0130] The above describes one or more embodiments of the response information generation method provided in this specification. Based on the same idea, this specification also provides a corresponding response information generation device, such as... Figure 6 As shown, the device includes:
[0131] The acquisition module 600 is used to acquire the user's input of the question to be answered;
[0132] The response information determination module 602 is used to input the question to be answered into the large language model, obtain the current response information output by the large language model, and obtain several historical response information output by the large language model based on the question to be answered;
[0133] The comparison information determination module 604 is used to determine the reply information to be compared from a number of historical reply information;
[0134] The difference determination module 606 is used to determine the difference result of the current reply based on the current reply information and the reply information to be compared; and to determine the current reply information as historical reply information.
[0135] Iteration module 608 is used to input the current response difference result and the question to be responded to into the large language model, redetermine the current response information, until the preset stop iteration condition is met;
[0136] The display module 610 is used to obtain and display the current response information when iteration stops.
[0137] Optionally, the response information determination module is specifically used to obtain preset first prompt information; based on the preset first prompt information and the question to be responded to, perform resampling to obtain a first resampling result; input the first resampling result into a large language model to obtain the current response information output by the large language model.
[0138] Optionally, the device further includes:
[0139] The resampling module is used to resample the current response information to obtain the resampled current response information.
[0140] Optionally, the difference determination module 606 is specifically used to obtain several historical response difference results to obtain a historical difference set; determine a second prompt message based on the historical difference set; and compare the current response information with the response information to be compared according to the second prompt message to determine the current response difference result.
[0141] Optionally, the difference determination module 606 is specifically used to compare the current reply information with the reply information to be compared according to the second prompt information, and determine the feedback result; and determine the current reply difference result according to the feedback result and the historical difference set.
[0142] Optionally, the iteration module 608 is specifically used to determine a third prompt message based on the current response difference result; to resample based on the third prompt message and the question to be answered to obtain a second resampling result; and to input the second resampling result and the current response difference result into the large language model to redetermine the current response message.
[0143] Optionally, the iteration module 608 is specifically used to, when there is no difference between the current response information and the response to be compared, perform resampling based on the preset first prompt information and the question to be answered, to obtain a first resampling result; input the first resampling result into the large language model to redetermine the current response information.
[0144] Optionally, the device further includes:
[0145] The confidence determination module is used to acquire all response information generated during the iteration process to obtain a response information set, which includes historical response information, current response information, and response information to be compared; determine the response distribution result of the response information set; and determine and display the confidence level of the current response information when the iteration stops based on the response distribution result.
[0146] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The provided method for generating reply information.
[0147] This instruction manual also provides Figure 7 The diagram shows the structure of the electronic device. Figure 7 As shown, at the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above. Figure 1 The method for generating response information is described above. Of course, besides software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0148] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0149] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0150] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0151] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0152] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0153] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0154] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0155] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0156] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0157] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0158] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0159] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0160] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0161] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0162] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0163] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this application.
Claims
1. A method for generating response information, the method comprising: Get the user's input for the question to be answered; Input the question to be answered into the large language model to obtain the current response information output by the large language model; And obtain several historical response information output by the large language model based on the question to be answered; Identify the response information to be compared from a number of historical response messages; Based on the current response information and the response information to be compared, the difference result of the current response is determined. The difference result of the current response is used as the input of the large language model, so that the large language model focuses on the difference information between the current response information and the response information to be compared; and the current response information is determined as historical response information. The current response difference result and the question to be responded to are input into the large language model to redetermine the current response information until the preset stopping iteration condition is met. Retrieve and display the current response information when iteration stops.
2. The method as described in claim 1, wherein the question to be answered is input into a large language model to obtain the current response information output by the large language model, specifically includes: Get the preset first prompt message; Based on the preset first prompt information and the question to be answered, resampling is performed to obtain the first resampling result; The first resampling result is input into the large language model to obtain the current response information output by the large language model.
3. The method as described in claim 1, before determining the difference result of the current response based on the current response information and the response information to be compared, the method further includes: The current response information is resampled to obtain the resampled current response information.
4. The method as described in claim 1, wherein determining the current response difference result based on the current response information and the response information to be compared specifically includes: Obtain several historical response difference results to obtain a historical difference set; Based on the aforementioned set of historical differences, the second prompt message is determined; Based on the second prompt information, the current response information is compared with the response information to be compared, and the difference result of the current response is determined.
5. The method as described in claim 4, wherein the current response information is compared with the response information to be compared based on the second prompt information to determine the difference result of the current response, specifically includes: Based on the second prompt information, the current response information is compared with the response information to be compared, and the feedback result is determined; Based on the feedback results and the historical difference set, the current response difference result is determined.
6. The method as described in claim 1, wherein the current response difference result and the question to be responded to are input into the large language model to redetermine the current response information, specifically includes: Based on the current response discrepancy results, determine the third prompt message; Based on the third prompt information and the question to be answered, resampling is performed to obtain the second resampling result; The second resampling result and the current response difference result are input into the large language model to redetermine the current response information.
7. The method as described in claim 2, further comprising: redetermining the current response information, specifically including: When the current response information is no different from the response to be compared, resampling is performed based on the preset first prompt information and the question to be answered to obtain the first resampling result; The first resampling result is input into the large language model to redetermine the current response information.
8. The method of claim 1, further comprising: All response information generated during the iteration process is obtained to obtain a response information set, which includes historical response information, current response information, and response information to be compared. Determine the response distribution result of the response information set; Based on the response distribution results, determine the confidence level of the current response information when stopping iteration, and display it.
9. A response information generation device, the device comprising: The acquisition module is used to acquire the user's input questions that need to be answered; The response information determination module is used to input the question to be answered into the large language model, obtain the current response information output by the large language model, and acquire several historical response information output by the large language model based on the question to be answered. The comparison information determination module is used to determine the reply information to be compared from a number of historical reply information; The difference determination module is used to determine the difference result of the current response based on the current response information and the response information to be compared. The difference result of the current response is used as the input of the large language model, so that the large language model pays attention to the difference information between the current response information and the response information to be compared; and the current response information is determined as historical response information. The iteration module is used to input the current response difference result and the question to be responded to into the large language model, redetermine the current response information, until the preset stop iteration condition is met; The display module is used to obtain and display the current response information when iteration stops.
10. The apparatus of claim 9, wherein the response information determination module is specifically configured to acquire preset first prompt information; resample based on the preset first prompt information and the question to be responded to, to obtain a first resampling result; and input the first resampling result into a large language model to obtain the current response information output by the large language model.
11. The apparatus of claim 9, further comprising: The resampling module is used to resample the current response information to obtain the resampled current response information.
12. The apparatus of claim 9, wherein the difference determination module is specifically configured to acquire a plurality of historical response difference results to obtain a historical difference set; determine a second prompt message based on the historical difference set; and compare the current response information with the response information to be compared according to the second prompt message to determine the current response difference result.
13. The apparatus of claim 12, wherein the difference determination module is specifically configured to compare the current response information with the response information to be compared according to the second prompt information, and determine the feedback result; and determine the current response difference result according to the feedback result and the historical difference set.
14. The apparatus of claim 9, wherein the iterative module is specifically configured to: determine a third prompt message based on the current response difference result; perform resampling based on the third prompt message and the question to be answered to obtain a second resampling result; input the second resampling result and the current response difference result into the large language model to redetermine the current response message.
15. The apparatus of claim 10, wherein the iteration module is specifically configured to, when the current response information is indistinguishable from the response to be compared, perform resampling based on the preset first prompt information and the question to be answered, to obtain a first resampling result; input the first resampling result into a large language model to redetermine the current response information.
16. The apparatus of claim 9, further comprising: The confidence determination module is used to acquire all response information generated during the iteration process to obtain a response information set, which includes historical response information, current response information, and response information to be compared; determine the response distribution result of the response information set; and determine and display the confidence level of the current response information when the iteration stops based on the response distribution result.
17. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 8.
18. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any one of claims 1 to 8.