Question answering method and device, equipment and storage medium
By combining explicit and implicit chain reasoning methods in a large language model, the problems of high computational overhead and low efficiency of explicit chain reasoning are solved, and accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510733204.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Existing large-scale language models have high computational overhead and lengthy reasoning processes in the explicit chain-thinking reasoning process, resulting in reduced efficiency.
The explicit thinking chain and the implicit thinking chain are combined and applied in the large language model. The reasoning process text and the first reasoning vector are generated through the explicit thinking chain reasoning, and the second reasoning vector is generated by combining the implicit thinking chain reasoning, and finally the reply text is generated.
The explicit bracelet is used to ensure the accuracy of the model's question-answering of input questions, while improving the efficiency of question-answering.
Smart Images

Figure CN120632039A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a question-answering method, apparatus, device, and storage medium. Background Art
[0002] With the remarkable progress made by Large Language Models (LLMs) in various reasoning tasks, Chain-of-Thought (CoT), as an effective reasoning technique, has gradually become one of the key methods to enhance the reasoning capabilities of language models.
[0003] In related technologies, explicit chain thinking is applied to the problem reasoning process of a large language model, so that the large language model explicitly generates intermediate reasoning steps during the problem reasoning process and presents the intermediate reasoning steps in the form of natural language, thereby obtaining the answer to the question through step-by-step reasoning.
[0004] However, explicit chain thinking requires the generation of a series of intermediate reasoning steps, which has high computational overhead and a lengthy reasoning process, resulting in reduced efficiency of the model in performing large-scale and complex reasoning tasks. Summary of the Invention
[0005] The present application provides a question-answering method, apparatus, device, and storage medium. The technical solution is as follows:
[0006] In one aspect, an embodiment of the present application provides a question-answering method, the method comprising:
[0007] Performing explicit thought chain reasoning on the input question through the first large language model to obtain a reasoning process text of the input question and a first reasoning vector, wherein the first reasoning vector is used to represent a result based on the explicit thought chain reasoning;
[0008] Performing implicit thought chain reasoning on the input question using the first language model to obtain a second reasoning vector for the input question, where the second reasoning vector is used to represent a result based on the implicit thought chain reasoning;
[0009] Based on the first inference vector and the second inference vector, generating a response text for the input question using the first large language model;
[0010] outputting the reasoning process text and the answer text;
[0011] Among them, the first language model is obtained based on the distillation learning of the second language model, and the second language model adopts the reasoning method based on the explicit thinking chain.
[0012] On the other hand, an embodiment of the present application provides a question-answering device, the device comprising:
[0013] a first reasoning module, configured to perform explicit chain of thought reasoning on an input question using a first large language model, to obtain a reasoning process text of the input question and a first reasoning vector, wherein the first reasoning vector is used to represent a result based on the explicit chain of thought reasoning;
[0014] a second reasoning module, configured to perform implicit thought chain reasoning on the input question using the first large language model to obtain a second reasoning vector for the input question, wherein the second reasoning vector is used to represent a result based on the implicit thought chain reasoning;
[0015] a text generation module, configured to generate a response text to the input question using the first large language model based on the first inference vector and the second inference vector;
[0016] A text output module, configured to output the reasoning process text and the response text;
[0017] Among them, the first language model is obtained based on the distillation learning of the second language model, and the second language model adopts the reasoning method based on the explicit thinking chain.
[0018] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the question-answering method described in the above aspect.
[0019] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the question-answering method described in the above aspects.
[0020] In another aspect, an embodiment of the present application provides a computer program product, comprising at least one instruction stored in a computer-readable storage medium. A processor of a computer device reads the at least one instruction from the computer-readable storage medium and executes the at least one instruction, causing the computer device to perform the question-and-answer method described in the above aspects.
[0021] In an embodiment of the present application, the explicit thought chain and the implicit thought chain are combined and applied in the first language model, so that in the process of executing question reasoning, the first language model performs explicit thought chain reasoning on the input question, and the reasoning process text and the first reasoning vector of the input question can be obtained; the first language model performs implicit thought chain reasoning on the input question, and the second reasoning vector of the input question can be obtained. Finally, based on the first reasoning vector and the second reasoning vector, the answer text of the input question can be obtained. The solution provided by the embodiment of the present application can ensure the accuracy of the first language model's question and answer of the input question based on the explicit thought chain, and at the same time improve the efficiency of the first language model's question and answer of the input question based on the implicit thought chain. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 A structural block diagram of a computer system provided by an exemplary embodiment of the present application is shown;
[0024] Figure 2 A flowchart of a question-answering method provided by an exemplary embodiment of the present application is shown;
[0025] Figure 3 A schematic diagram showing a first large language model performing question reasoning provided by an exemplary embodiment of the present application is shown;
[0026] Figure 4 A schematic diagram of a question-and-answer interface provided by an exemplary embodiment of the present application is shown;
[0027] Figure 5 A flowchart of training a first language model provided by an exemplary embodiment of the present application is shown;
[0028] Figure 6 A schematic diagram of training a first language model provided by an exemplary embodiment of the present application is shown;
[0029] Figure 7 A flowchart of a question-answering method provided by another exemplary embodiment of the present application is shown;
[0030] Figure 8 A flowchart of a question-answering method provided by another exemplary embodiment of the present application is shown;
[0031] Figure 9A structural block diagram of a question-answering device provided by an exemplary embodiment of the present application is shown;
[0032] Figure 10 A schematic structural diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0035] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0036] It should be understood that although the terms first, second, etc. may be used in this application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0037] First, a brief introduction to the terms involved in the embodiments of this application is given:
[0038] Large Language Model (LLM): A neural network model trained on large amounts of text data. It learns the patterns, structure, and semantics of language, enabling the understanding and generation of natural language. Large Language Models can be used for a variety of natural language processing tasks, including but not limited to text generation, translation, question answering, text classification, and sentiment analysis.
[0039] Explicit Chain of Thought (Explicit CoT): A method that explicitly and step-by-step demonstrates the thought process and intermediate reasoning steps of a natural language processing model when solving complex problems or reasoning. By applying explicit CoT, the model's decision-making process becomes transparent and explainable.
[0040] Implicit Chain of Thought (CoT): When a natural language processing model solves complex problems or performs reasoning, the reasoning process within the model is not explicitly displayed but hidden within the model's internal mechanisms. The model directly provides the final answer without providing the intermediate reasoning steps and thought processes.
[0041] The Attention Mechanism is a computational model that mimics human visual attention. It aims to dynamically focus on the most relevant parts of input data, rather than treating all input information equally. The core components of the attention mechanism are Q (Query), K (Key), and V (Value). These components implement a weighted approach to input data through a specific computational process, enabling the model to dynamically focus on the information most relevant to the task at hand.
[0042] Please refer to Figure 1 , which shows a structural block diagram of a computer system provided by an exemplary embodiment of the present application, wherein the computer system may include a terminal 110 and a server 120. The terminal 110 and the server 120 perform data communication via a communication network. Optionally, the communication network may be a wired network or a wireless network, and the communication network may be at least one of a local area network, a metropolitan area network, and a wide area network.
[0043] Terminal 110 is an electronic device that has an application with a question-and-answer function installed. The question-and-answer function can be a function of a native application on terminal 110 or a function of a third-party application. Terminal 110 can be a smartphone, tablet computer, laptop computer, desktop computer, smart TV, wearable device, or in-vehicle terminal, etc. Figure 1 The example of the terminal 110 being a desktop computer is merely used for illustration, but is not intended to be limiting.
[0044] The server 120 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. In the embodiment of the present application, the server 120 may be a background server of an application program having a question-answering function.
[0045] In some embodiments, there is data interaction between the server 120 and the terminal 110. Schematically, as Figure 1As shown, after receiving a question input operation, terminal 110 sends the received input question to server 120. Server 120 then performs explicit chain of thought reasoning on the input question using the first large language model to obtain the reasoning process text and the first reasoning vector of the input question. Furthermore, server 120 performs implicit chain of thought reasoning on the input question using the first large language model to obtain the second reasoning vector of the input question. Furthermore, based on the first and second reasoning vectors, server 120 generates a reply text for the input question using the first large language model, ultimately outputting the reasoning process text and the reply text. These are then returned to terminal 110 so that terminal 110 can display the reasoning process text and the reply text to the user.
[0046] In combination with the above introduction, the question-answering method provided in this application is explained. The method can be executed by a server or a terminal, or by a server and a terminal together.
[0047] Please refer to Figure 2 , which shows a flowchart of a question-answering method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method applied to a computer device (including a terminal and / or a server) as an example. The method includes the following steps:
[0048] Step 201 : Perform explicit chain of thought reasoning on the input question through the first large language model to obtain the reasoning process text of the input question and a first reasoning vector, where the first reasoning vector is used to represent the result based on the explicit chain of thought reasoning.
[0049] In some embodiments, in order to enable the first language model to have the ability of explicit thought chain reasoning and implicit thought chain reasoning, the computer device can obtain the first language model based on distillation learning of the second language model, wherein the second language model adopts a reasoning method based on explicit thought chain.
[0050] In an embodiment of the present application, explicit thought chains and implicit thought chains are combined and applied to the first language model, so that in the process of performing question reasoning, the first language model can respectively perform explicit thought chain reasoning and implicit thought chain reasoning on the input question. Compared with applying the second language model to perform only explicit thought chain reasoning on the input question, applying the first language model proposed in this application to perform question reasoning can achieve the application of explicit thought chain reasoning to ensure the accuracy of the model's question-answering of the input question, while applying implicit thought chain reasoning to improve the model's question-answering efficiency for the input question.
[0051] Optionally, the first language model has explicit thought chain reasoning capabilities and implicit thought chain reasoning capabilities. The first language model can be applied to a variety of scenarios that require reasoning, such as natural language processing, mathematical reasoning, logical reasoning, artificial intelligence reasoning tasks, etc., especially tasks involving complex reasoning chains.
[0052] Optionally, after the question is input into the first language model, the first language model may perform explicit thought chain reasoning and implicit thought chain reasoning on the input question in parallel.
[0053] Optionally, explicit chaining means that when the first language model performs problem reasoning, its thought process and intermediate reasoning steps can be displayed in an explicit, step-by-step manner, making the reasoning process explainable and transparent. Optionally, during explicit chaining reasoning, the first language model can generate a series of explicit reasoning step tokens. These tokens represent the intermediate explicit reasoning steps and intermediate results of the explicit chaining reasoning.
[0054] Alternatively, an explicit thought chain can be represented as {S1, S2, S3, …Sn}. Each reasoning step based on the explicit thought chain is explicitly generated and output.
[0055] In some embodiments, upon receiving a question input operation, the computer device performs explicit chaining reasoning on the input question using the first large language model, thereby obtaining a reasoning text for each explicit reasoning step of the input question, i.e., a reasoning process text. Furthermore, after completing the explicit chaining reasoning, the computer device can obtain a first reasoning vector for the input question.
[0056] The input question can be in the form of text, images, documents, website links, etc., which is not limited in the embodiments of the present application. Optionally, the input question is a statement used to obtain information or answers, and the input question has a clear questioning intention. For example, the input question is "What are the specific functions of product A?"
[0057] Optionally, the reasoning process text includes the reasoning text corresponding to each explicit reasoning step. For example, the reasoning process text may be "First calculate A, then calculate B, and finally obtain C." Each explicit reasoning step corresponds to a link in the explicit thinking chain.
[0058] Optionally, the first reasoning vector is used to represent the result of explicit chaining reasoning. That is, in this embodiment of the application, after executing explicit chaining reasoning, the response text corresponding to the explicit chaining reasoning is not directly output. Instead, only the first reasoning vector is generated to facilitate the subsequent combination of the result of the explicit chaining reasoning with the result of the implicit chaining reasoning to obtain the final response text for the input question.
[0059] Step 202 : Perform implicit thought chain reasoning on the input question through the first large language model to obtain a second reasoning vector of the input question, where the second reasoning vector is used to represent a result based on the implicit thought chain reasoning.
[0060] Optionally, implicit thought chaining means that when the first language model performs question reasoning, the internal reasoning process of the model is not explicitly output, but is hidden within the model. In other words, the model does not output the intermediate reasoning steps and thinking process based on implicit thought chaining reasoning.
[0061] Optionally, the implicit thought chain can be represented as {Z1, Z2, Z3, … Zk}. Each reasoning step based on the implicit thought chain is hidden within the model and is not output by the model.
[0062] In some embodiments, upon receiving a question input operation, the computer device performs implicit thought chain reasoning on the input question through the first large language model, thereby obtaining a second reasoning vector of the input question.
[0063] The second reasoning vector is used to represent the result of implicit chain reasoning. Specifically, in this embodiment of the application, after executing implicit chain reasoning, the corresponding response text is not directly output. Instead, only the second reasoning vector is generated to facilitate the subsequent combination of the implicit chain reasoning result with the explicit chain reasoning result to obtain the final response text for the input question.
[0064] Step 203: Generate a response text to the input question using a first large language model based on the first inference vector and the second inference vector.
[0065] In some embodiments, after completing the explicit chain of thought reasoning and the implicit chain of thought reasoning respectively, the computer device can generate a response text for the input question through the first large language model based on the first reasoning vector corresponding to the explicit chain of thought reasoning and the second reasoning vector corresponding to the implicit chain of thought reasoning.
[0066] Optionally, an attention mechanism may be applied to the first language model to fuse the first inference vector and the second inference vector to obtain a response text to the input question.
[0067] Indicative, such as Figure 3 As shown, the computer device inputs the input question 301 into the first language model, so that the first language model performs explicit thought chain reasoning and implicit thought chain reasoning on the input question 301, that is, a question and answer result 302 corresponding to the input question 301 can be obtained. The question and answer result 302 includes both the reasoning process text obtained based on the explicit thought chain reasoning and the reply text obtained based on the explicit thought chain reasoning and the implicit thought chain reasoning.
[0068] Among them, in the process of explicit thought chain reasoning, the first language model generates the reasoning text 303 of each explicit reasoning step through discrete tokens, thereby realizing the display of the reasoning process of explicit thought chain reasoning to the user; and in the process of implicit thought chain reasoning, the first language model uses continuous token hidden layers to represent each implicit reasoning step, thereby realizing efficient problem reasoning based on continuous hidden layer reasoning representation 304.
[0069] Step 204: Output the reasoning process text and the response text.
[0070] Finally, after obtaining the reasoning process text and the reply text, the computer device can output the reasoning process text and the reply text, so that the user can obtain the answer result of the input question and understand the explicit reasoning steps of the first language model from the reasoning process text.
[0071] Indicative, such as Figure 4 As shown, the first language model proposed in this application is applied to the question-answering interface, so that users can input questions in the question-answering interface. When the computer device receives the input question, it performs explicit thinking chain reasoning and implicit thinking chain reasoning on the input question through the first language model, thereby obtaining the reasoning process text and the answer text of the input question, and displays the reasoning process text and the answer text in the result display box 401 of the question-answering interface.
[0072] In summary, in the embodiment of the present application, the explicit thought chain and the implicit thought chain are combined and applied in the first language model, so that in the process of executing question reasoning, the first language model performs explicit thought chain reasoning on the input question, and the reasoning process text and the first reasoning vector of the input question can be obtained; the first language model performs implicit thought chain reasoning on the input question, and the second reasoning vector of the input question can be obtained. Finally, the answer text of the input question can be obtained based on the first reasoning vector and the second reasoning vector. The solution provided in the embodiment of the present application can ensure the accuracy of the first language model's question and answer of the input question based on the explicit thought chain, and at the same time improve the efficiency of the first language model's question and answer of the input question based on the implicit thought chain.
[0073] In some embodiments, in order to improve the question-answering accuracy of the first language model while improving the efficiency of question-answering reasoning, a discriminant model can also be introduced in the explicit thinking chain reasoning process to discriminate the accuracy of each explicit reasoning step in the explicit thinking chain reasoning through the discriminant model.
[0074] Optionally, the discriminant model can predict whether the explicit reasoning steps are accurate based on the input reasoning texts of each explicit reasoning step, thereby obtaining the reasoning discrimination results corresponding to the explicit reasoning steps.
[0075] In one possible implementation, in the process of performing explicit thought chain reasoning on the input question through the first large language model, the computer device discriminates each explicit reasoning step through the discriminant model, thereby obtaining the reasoning judgment result of each explicit reasoning step, and based on the reasoning judgment result of each explicit reasoning step, generates the reasoning process text of the input question and the first reasoning vector through the first large language model.
[0076] Optionally, the reasoning process text may include reasoning text corresponding to n explicit reasoning steps, where n is a positive integer.
[0077] Optionally, the discriminant model can perform a classification task on the inference text corresponding to each explicit reasoning step, thereby using the classification result of the classification task as the inference discrimination result corresponding to the explicit reasoning step. For example, the discriminant model performs a two-class classification task and outputs a classification result of "+1" if the explicit reasoning step is accurate; and a classification result of "-1" if the explicit reasoning step is incorrect. For another example, the discriminant model can also perform a three-class classification task and output a classification result of "+1" if the explicit reasoning step is accurate; a classification result of "-1" if the explicit reasoning step is incorrect; and a classification result of "0" if the explicit reasoning step has no significant effect.
[0078] In one possible implementation, the computer device performs the i-th step of explicit thought chain reasoning on the input question through the first large language model to obtain the i-th step reasoning text corresponding to the i-th explicit reasoning step, where i is a positive integer and i is less than or equal to n. Furthermore, in order to avoid affecting the accuracy of subsequent explicit reasoning steps when the i-th explicit reasoning step is wrong, the computer device needs to promptly judge the accuracy of the i-th explicit reasoning step through the discriminant model. Optionally, the computer device judges the i-th step reasoning text through the discriminant model to obtain the i-th reasoning judgment result of the i-th explicit reasoning step, and the i-th reasoning judgment result is used to characterize whether the i-th explicit reasoning step is accurate.
[0079] Optionally, when the i-th reasoning judgment result indicates that the i-th explicit reasoning step is accurate, the computer device can perform the i+1-th explicit thinking chain reasoning on the input question through the first largest language model to obtain the i+1-th reasoning text of the input question, and then continue to judge the i+1-th reasoning text through the discrimination model, and so on, until the accuracy judgment of each explicit reasoning step is completed.
[0080] Optionally, when the i-th reasoning judgment result represents that the i-th explicit reasoning step is inaccurate, the computer device needs to re-execute the i-th step explicit thinking chain reasoning on the input problem through the first largest language model to obtain the i-th step reasoning text of the input problem, and re-judge the i-th step reasoning text through the discrimination model to obtain a new i-th reasoning judgment result, and then, when the current i-th reasoning judgment result represents that the i-th explicit reasoning step is accurate, continue to execute the i+1-th step explicit thinking chain reasoning through the first largest language model, and so on, until the accuracy judgment of each explicit reasoning step is completed.
[0081] Finally, when the n-1th reasoning judgment result represents that the n-1th explicit reasoning step is accurate, the computer device can perform the nth step of explicit thinking chain reasoning on the input question through the first large language model, thereby obtaining the nth step reasoning text of the input question and the first reasoning vector.
[0082] Optionally, after obtaining the n-th step reasoning text, the computer device can continue to judge the n-th explicit reasoning step through the discriminant model, and when the n-th reasoning judgment result represents that the n-th explicit reasoning step is accurate, the above-mentioned n-th step reasoning text and the first reasoning vector are determined as the final explicit thinking chain reasoning result; when the n-th reasoning judgment result represents that the n-th explicit reasoning step is inaccurate, it is necessary to re-execute the n-th step explicit thinking chain reasoning to obtain a new n-step reasoning text and the first reasoning vector.
[0083] In the above embodiment, by introducing a discriminant model to judge the accuracy of each explicit reasoning step during the process of executing explicit thought chain reasoning on the first language model, the problem of an inaccurate explicit reasoning step affecting the accuracy of subsequent explicit reasoning steps can be avoided, thereby improving the overall accuracy of the explicit thought chain reasoning.
[0084] Moreover, by using the discriminant model to discriminate each explicit reasoning step, the first language model can dynamically select the optimal reasoning path during the problem reasoning process, which is beneficial for the first language model to handle more types of complex reasoning tasks.
[0085] Regarding the process of generating the reply text, in some embodiments, after obtaining the first inference vector and the second inference vector, the computer device can apply the attention mechanism through the first large language model to fuse the first inference vector and the second inference vector to obtain the reply text of the input question.
[0086] In one possible implementation, the computer device first performs a linear transformation on the first inference vector and the second inference vector through the first large language model to obtain a query vector, a key vector, and a value vector for attention calculation, and then performs attention calculation on the query vector, the key vector, and the value vector through the first large language model to obtain attention weights, and finally performs weighted fusion on the attention weights and the value vector through the first large language model to generate a response text for the input question.
[0087] Optionally, during the attention calculation process, when focusing on the attention of the first inference vector to the second inference vector, the dot product of the query vector of the first inference vector and the key vector of the second inference vector can be calculated and divided by a scaling factor to obtain an attention score; when focusing on the attention of the second inference vector to the first inference vector, the dot product of the query vector of the second inference vector and the key vector of the first inference vector is calculated and divided by a scaling factor to obtain an attention score. Furthermore, by normalizing the attention scores, the attention weights can be obtained.
[0088] By fusing the first reasoning vector and the second reasoning vector based on the attention mechanism, it is possible to integrate the results of explicit thought chain reasoning and implicit thought chain reasoning, thereby improving the comprehensiveness and accuracy of the reply text.
[0089] In some embodiments, to combine explicit and implicit chaining reasoning in the first language model, it is also necessary to optimize the reasoning capabilities of the first language model based on distillation learning of the second language model. The following describes the process of training the first language model through specific examples.
[0090] Please refer to Figure 5 , which shows a flowchart of training the first language model provided by an exemplary embodiment of the present application. This embodiment is described by taking the method used in a computer device (including a terminal and / or a server) as an example. The method includes the following steps:
[0091] Step 501: Perform explicit thought chain reasoning and implicit thought chain reasoning on a sample question through a first large language model to obtain a first sample reasoning process text and a first sample answer text of the sample question.
[0092] Optionally, in order to ensure that the first language model has the ability to perform reasoning on questions of any form, the sample questions can be in the form of text, pictures, documents, website links, etc., which is not limited in this embodiment of the present application.
[0093] In one possible implementation, a computer device inputs a sample question into a first large language model, and performs explicit thought chain reasoning and implicit thought chain reasoning on the sample question through the first large language model, thereby obtaining a first sample reasoning process text and a first sample answer text of the sample question.
[0094] Optionally, the computer device performs explicit thought chain reasoning on the sample question through the first large language model to obtain a first sample reasoning process text and a first sample reasoning vector of the sample question, and performs implicit thought chain reasoning on the sample question through the first large language model to obtain a second sample reasoning vector of the sample question, thereby generating a first sample response text of the sample question through the first large language model based on the first sample reasoning vector and the second sample reasoning vector.
[0095] In one possible implementation, during the training process of the first language model, the computer device may also apply the discriminant model to perform discrimination on the explicit thought chain reasoning process, thereby using the sample reasoning discrimination result as a reward signal to improve the accuracy of the training reasoning process of the first language model.
[0096] Optionally, during the process of performing explicit chain reasoning on the sample question using the first large language model, the computer device may use the discriminant model to discriminate each sample explicit reasoning step, and based on the reasoning and discrimination results of each sample explicit reasoning step, generate a first sample reasoning process text and a first sample reasoning vector for the sample question using the first large language model. The specific implementation of introducing the discriminant model during training can be referenced in the above-mentioned application-side embodiment and will not be elaborated on in this embodiment.
[0097] By introducing a discriminant model during the model training process, the sample reasoning discrimination results of each sample explicit reasoning step can be used as reward signals to guide the first language model to select the optimal reasoning path and adjust the reasoning strategy of each step in real time. This enables the first language model to dynamically adjust the generation method of the explicit thinking chain according to the complexity and specific needs of the question-answering task, thereby improving the model's reasoning ability.
[0098] Step 502: Execute explicit thought chain reasoning on the sample question through the second largest language model to obtain a second sample reasoning process text and a second sample answer text of the sample question.
[0099] Optionally, the second language model is a pre-trained large language model. Optionally, the second language model shares the same neural network architecture as the first language model, but the second language model is tasked with learning explicit chaining reasoning, while the first language model is tasked with learning explicit chaining reasoning and implicit chaining reasoning.
[0100] In one possible implementation, the computer device inputs the sample question into the second largest language model, and performs explicit thought chain reasoning on the sample question through the second largest language model, thereby obtaining a second sample reasoning process text and a second sample answer text of the sample question.
[0101] Step 503: Train a first large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, the second sample reply text, and the true value of the answer to the sample question.
[0102] After obtaining the first sample reasoning process text and the first sample answer text of the sample question generated by the first largest language model, and the second sample reasoning process text and the second sample answer text of the sample question generated by the second largest language model, the computer device can train the first largest language model based on the distillation learning of the second largest language model.
[0103] In one possible implementation, the computer device determines the inference loss of the first large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, the second sample reply text, and the true value of the answer to the sample question, thereby training the first large language model.
[0104] In the above embodiment, the second largest language model with explicit thought chain reasoning ability is used as the teacher model. Through distillation learning, the first largest language model can learn the problem reasoning ability of the second largest language model, thereby ensuring the problem reasoning accuracy of the first largest language model.
[0105] Optionally, during the distillation learning process, the second largest language model can be frozen and the first largest language model can be trained separately, or the first largest language model and the second largest language model can be trained jointly.
[0106] In some embodiments, when training the first large language model alone, the computer device may determine the distillation loss between the first large language model and the second large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text, and determine the first inference loss of the first large language model based on the first sample reply text and the true value of the answer to the sample question, thereby training the first large language model based on the distillation loss and the first inference loss.
[0107] Among them, regarding the method of determining the first reasoning loss, the computer device can adopt the cross-entropy loss calculation method. In the process of generating the first sample reply text, the cross-entropy loss is calculated for the explicit reasoning step and the implicit reasoning step respectively, so as to obtain the first reasoning loss.
[0108] Alternatively, the loss function during explicit thought chain reasoning can be expressed as Where n represents the number of explicit reasoning steps, that is, the length of the explicit thinking chain, P(S i |Q) represents the i-th step reasoning S generated in the explicit thinking chain reasoning i The probability of , Q represents the sample problem.
[0109] Alternatively, the loss function in the implicit thought chain reasoning process can be expressed as Among them, k represents the number of implicit reasoning steps, that is, the length of the implicit thinking chain, P(Z i |Q) represents the i-th step reasoning Z generated in the explicit thinking chain reasoning i The probability of , Q represents the sample problem.
[0110] Optionally, the computer device may perform weighted fusion of the explicit thinking chain reasoning loss and the implicit thinking chain reasoning loss to obtain the first reasoning loss L Student .
[0111] Among them, regarding the method of determining the distillation loss, the L1 distance loss calculation method can be used to determine the distillation loss between the first largest language model and the second largest language model.
[0112] In one possible implementation, the computer device may determine the first hidden activation value of each hidden layer in the first large language model during the process of the first large language model outputting the first sample reasoning process text and the first sample reply text, and determine the second hidden activation value of each hidden layer in the second large language model during the process of the second large language model outputting the second sample reasoning process text and the second sample reply text, thereby aligning the first hidden activation value of each hidden layer in the first large language model with the second hidden activation value of each hidden layer in the second large language model, that is, determining the distillation loss based on the difference between the aligned first hidden activation value and the second hidden activation value.
[0113] Regarding the alignment method of the first hidden activation value and the second hidden activation value, optionally, when the neural network architecture of the first large language model and the neural network architecture of the second large language model are the same, each hidden layer in the first large language model corresponds one-to-one with each hidden layer in the second large language model; when the neural network architecture of the first large language model and the neural network architecture of the second large language model are different, it is necessary to align the first hidden activation value and the second hidden activation value based on the number of hidden layers in the first large language model and the number of hidden layers in the second large language model. For example, if the number of hidden layers in the first large language model is 3 and the number of hidden layers in the second large language model is 5, the first hidden layer in the first large language model can be aligned with the first hidden layer in the second large language model, the second hidden layer in the first large language model can be aligned with the third hidden layer in the second large language model, and the third hidden layer in the first large language model can be aligned with the fifth hidden layer in the second large language model.
[0114] Alternatively, the distillation loss can be expressed as Where L represents the number of hidden layers in the first language model, represents the first hidden activation value of the lth hidden layer in the first language model, represents the second hidden activation value of the hidden layer in the second largest language model that is aligned with the lth hidden layer in the first largest language model.
[0115] In another possible implementation, to improve the computational efficiency of distillation loss, the computer device can directly determine the distillation loss based on the last hidden activation value before the model output. Optionally, the computer device can determine a third hidden activation value in the first large language model during the process of the first large language model outputting a first sample reasoning process text and a first sample response text. This third hidden activation value is the last hidden activation value before the first large language model outputs the first sample response text. Furthermore, the computer device can determine a fourth hidden activation value in the second large language model during the process of the second large language model outputting a second sample reasoning process text and a second sample response text. This fourth hidden activation value is the last hidden activation value before the second large language model outputs the second sample response text. Furthermore, the distillation loss is determined based on the difference between the third and fourth hidden activation values.
[0116] Finally, we get the first inference loss L student and distillation loss L KD After that, the computer device can perform weighted fusion on the first inference loss and the distillation loss to obtain a weighted loss, thereby training the first language model based on the weighted loss.
[0117] In other embodiments, when jointly training the first large language model and the second large language model, the computer device may determine the distillation loss between the first large language model and the second large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text, and determine the first inference loss of the first large language model based on the first sample reply text and the true value of the answer to the sample question; and determine the second inference loss of the second large language model based on the second sample reply text and the true value of the answer to the sample question, thereby jointly training the first large language model and the second large language model based on the distillation loss, the first inference loss, and the second inference loss.
[0118] The above-described embodiments may be referred to for determining the first inference loss and the distillation loss. Regarding determining the second inference loss, the computer device may also employ a cross-entropy loss calculation method. During the process of generating the second sample reply text, the cross-entropy loss calculation is performed for the explicit inference step to obtain the second inference loss.
[0119] Optionally, the explicit thought chain in the second language model can be expressed as {T1, T2, T3, ... Tm}, and the loss function of the second inference loss can be expressed as Where m represents the number of explicit reasoning steps in the second largest language model, that is, the length of the explicit thinking chain, P(T i |Q) represents the i-th step of reasoning T generated in the explicit thinking chain reasoning i The probability of , Q represents the sample problem.
[0120] Finally, we get the first inference loss L student , the second inference loss L teacher and distillation loss L KD Afterwards, the computer device can fuse the distillation loss, the first inference loss and the second inference loss based on the first loss weight, the second loss weight and the third loss weight to obtain the comprehensive inference loss, thereby jointly training the first language model and the second language model based on the comprehensive inference loss.
[0121] Alternatively, the comprehensive inference loss can be expressed as L = αL teacher +βL student +γL KD , where γ, β, and α are the first loss weight, the second loss weight, and the third loss weight, respectively, which are used to balance the distillation loss, the first inference loss, and the second inference loss.
[0122] Indicative, such as Figure 6As shown, the first large language model 601 serves as a student model and the second large language model 602 serves as a teacher model. The computer device performs question reasoning on the sample question through the first large language model 601 and the second large language model 602 respectively, thereby obtaining a first sample reasoning process text and a first sample answer text of the sample question, as well as a second sample reasoning process text and a second sample answer text of the sample question. The computer device determines a first reasoning loss based on the first sample answer text and the true value of the answer, determines a second reasoning loss based on the second sample answer text and the true value of the answer, and determines a distillation loss based on the hidden activation values of each hidden layer in the first large language model 601 and the hidden activation values of each hidden layer in the second large language model 602. The first large language model 601 is trained by combining the first reasoning loss, the second reasoning loss and the distillation loss.
[0123] It should be noted that when jointly training the first and second language models, since the second language model has explicit thought chain reasoning capabilities, the first language model needs to learn both explicit and implicit thought chain reasoning capabilities. Therefore, in order to reduce resources and computing costs, computer equipment can also use self-distillation to train the first language model.
[0124] Optionally, the part of the first language model that performs explicit chain of thought reasoning can be directly inherited from the second language model, and during the training process, the network layer that performs explicit chain of thought reasoning can guide the network layer that performs implicit chain of thought reasoning, so that the first language model can optimize the explicit chain of thought reasoning ability while learning the implicit chain of thought reasoning ability.
[0125] In one possible implementation, when the first language model is trained using self-distillation, the student task loss can be calculated based on the implicit chain reasoning process, the teacher task loss can be calculated based on the explicit chain reasoning process, and the distillation loss can be calculated based on the difference between the implicit chain reasoning and the explicit chain reasoning. Optionally, the computer device can calculate the distillation loss by L1 distance loss based on the difference between the last hidden activation value before the second sample inference vector is generated by the implicit chain reasoning and the last hidden activation value before the first sample inference vector is generated by the explicit chain reasoning.
[0126] In the above embodiment, in the process of training the first language model, knowledge distillation is performed by minimizing the hidden activation difference between the first language model and the second language model, so that the first language model can not only learn the latent reasoning representation generated by implicit thought chain reasoning, but also learn the structure and logic of the reasoning process from explicit thought chain reasoning, thereby balancing the explicit thought chain reasoning ability and implicit thought chain reasoning ability of the first language model, so that the first language model can ensure the transparency and accuracy of the reasoning process while maintaining the efficiency of question and answer.
[0127] Moreover, while the first language model learns the implicit thought chain reasoning ability, it can also use the reasoning vector obtained based on the implicit thought chain reasoning to assist the explicit thought chain reasoning, thereby reducing the explicit reasoning steps of the first language model in the explicit thought chain reasoning process, shortening the reasoning time, avoiding the lengthy reasoning steps in the second language model, and reducing the dependence on natural language tags and computational overhead.
[0128] Applications in AI question answering applications:
[0129] In one application scenario, the question-answering method provided in the embodiment of the present application can be applied to an AI (artificial intelligence) question-answering application. Figure 7 , which shows a flowchart of a question-answering method provided by another exemplary embodiment of the present application. This embodiment is described by taking the method applied to a computer device (including a terminal and / or a server) as an example. The method includes the following steps:
[0130] Step 701 , performing explicit thought chain reasoning on the input question through the first large language model to obtain the reasoning process text of the input question and a first reasoning vector, where the first reasoning vector is used to represent the result based on the explicit thought chain reasoning.
[0131] Optionally, in an AI question-and-answer application, a question-and-answer interface is displayed, and in response to a user input operation, the computer device determines the input question. Furthermore, the computer device performs explicit chain-of-thought reasoning on the input question using the first large language model, thereby obtaining reasoning text for each explicit reasoning step of the input question, i.e., the reasoning process text. Furthermore, after completing the explicit chain-of-thought reasoning, the computer device can obtain a first reasoning vector for the input question.
[0132] Step 702: Perform implicit thought chain reasoning on the input question through the first large language model to obtain a second reasoning vector of the input question, where the second reasoning vector is used to represent a result based on the implicit thought chain reasoning.
[0133] Optionally, in the AI question-answering application, a question-answering interface is displayed, and in response to the user's question input operation, the computer device determines the input question. Then, the computer device performs implicit thought chain reasoning on the input question using the first large language model, thereby obtaining a second reasoning vector for the input question.
[0134] Step 703: Generate a response text to the input question using a first large language model based on the first inference vector and the second inference vector.
[0135] In one possible implementation, the computer device first performs a linear transformation on the first inference vector and the second inference vector through the first large language model to obtain a query vector, a key vector, and a value vector for attention calculation, and then performs attention calculation on the query vector, the key vector, and the value vector through the first large language model to obtain attention weights, and finally performs weighted fusion on the attention weights and the value vector through the first large language model to generate a response text for the input question.
[0136] Step 704: Display the reasoning process text and the answer text in the question and answer application interface.
[0137] After obtaining the reasoning process text and the answer text of the input question, the computer device can display the reasoning process text and the answer text in the question and answer application interface.
[0138] Applied to AI question answering web pages:
[0139] In one application scenario, the question-answering method provided in the embodiment of the present application can be applied to an AI question-answering webpage. Figure 8 , which shows a flowchart of a question-answering method provided by another exemplary embodiment of the present application. This embodiment is described by taking the method applied to a computer device (including a terminal and / or a server) as an example. The method includes the following steps:
[0140] Step 801 , performing explicit thought chain reasoning on an input question through a first large language model, obtaining a reasoning process text of the input question and a first reasoning vector, where the first reasoning vector is used to represent a result based on the explicit thought chain reasoning.
[0141] Optionally, a question-and-answer interface is displayed on the AI Q&A webpage. In response to a user input operation, the computer device determines the input question. Furthermore, the computer device performs explicit chain reasoning on the input question using the first language model, thereby obtaining reasoning text for each explicit reasoning step of the input question, i.e., the reasoning process text. Furthermore, after completing the explicit chain reasoning, the computer device can obtain a first reasoning vector for the input question.
[0142] Step 802: Perform implicit thought chain reasoning on the input question through the first large language model to obtain a second reasoning vector of the input question, where the second reasoning vector is used to represent the result based on the implicit thought chain reasoning.
[0143] Optionally, a question-and-answer interface is displayed on the AI Q&A webpage, and in response to the user's question input operation, the computer device determines the input question. Furthermore, the computer device performs implicit thought chain reasoning on the input question using the first large language model, thereby obtaining a second reasoning vector for the input question.
[0144] Step 803: Generate a response text to the input question using a first large language model based on the first inference vector and the second inference vector.
[0145] In one possible implementation, the computer device first performs a linear transformation on the first inference vector and the second inference vector through the first large language model to obtain a query vector, a key vector, and a value vector for attention calculation, and then performs attention calculation on the query vector, the key vector, and the value vector through the first large language model to obtain attention weights, and finally performs weighted fusion on the attention weights and the value vector through the first large language model to generate a response text for the input question.
[0146] Step 804: Display the reasoning process text and the answer text on the question-and-answer webpage.
[0147] After obtaining the reasoning process text and the answer text of the input question, the computer device can display the reasoning process text and the answer text on the question and answer webpage.
[0148] Please refer to Figure 9 , which shows a structural block diagram of a question-answering device provided by an exemplary embodiment of the present application, the device comprising:
[0149] A first reasoning module 901 is configured to perform explicit chain of thought reasoning on an input question using a first large language model to obtain a reasoning process text of the input question and a first reasoning vector, wherein the first reasoning vector is used to represent a result of the explicit chain of thought reasoning;
[0150] A second reasoning module 902 is configured to perform implicit thought chain reasoning on the input question using the first large language model to obtain a second reasoning vector for the input question, where the second reasoning vector is used to represent a result based on the implicit thought chain reasoning;
[0151] A text generation module 903 is configured to generate a response text to the input question using the first language model based on the first inference vector and the second inference vector;
[0152] A text output module 904, configured to output the reasoning process text and the response text;
[0153] Among them, the first language model is obtained based on the distillation learning of the second language model, and the second language model adopts the reasoning method based on the explicit thinking chain.
[0154] Optionally, the first reasoning module 901 is configured to:
[0155] In the process of performing explicit thought chain reasoning on the input question by using the first language model, each explicit reasoning step is discriminated by using the discriminant model to obtain a reasoning and discrimination result of each explicit reasoning step;
[0156] Based on the reasoning judgment results of the respective explicit reasoning steps, a reasoning process text of the input question and the first reasoning vector are generated through the first large language model.
[0157] Optionally, the reasoning process text includes reasoning text corresponding to n explicit reasoning steps, where n is a positive integer;
[0158] The first reasoning module 901 is configured to:
[0159] Performing an i-th step of explicit thought chain reasoning on the input question using the first language model to obtain an i-th step of reasoning text corresponding to the i-th explicit reasoning step, where i is a positive integer and is less than or equal to n;
[0160] The i-th reasoning text is judged by a discriminant model to obtain an i-th reasoning judgment result of the i-th explicit reasoning step, and the i-th reasoning judgment result is used to indicate whether the i-th explicit reasoning step is accurate.
[0161] Optionally, the first reasoning module 901 is configured to:
[0162] When the i-th reasoning judgment result indicates that the i-th explicit reasoning step is accurate, performing the i+1-th explicit thought chain reasoning on the input question through the first large language model to obtain the i+1-th reasoning text of the input question;
[0163] When the i-th reasoning judgment result indicates that the i-th explicit reasoning step is inaccurate, re-execute the i-th explicit thought chain reasoning on the input question using the first large language model to obtain the i-th reasoning text of the input question;
[0164] When the n-1th reasoning judgment result indicates that the n-1th explicit reasoning step is accurate, the nth step of explicit thinking chain reasoning is performed on the input question through the first large language model to obtain the nth step reasoning text of the input question and the first reasoning vector.
[0165] Optionally, the text generation module 903 is used to:
[0166] Performing a linear transformation on the first inference vector and the second inference vector using the first large language model to obtain a query vector, a key vector, and a value vector;
[0167] performing an attention calculation on the query vector, the key vector, and the value vector using the first language model to obtain an attention weight;
[0168] The first language model performs weighted fusion on the attention weight and the value vector to generate the answer text of the input question.
[0169] Optionally, the device further includes:
[0170] A first sample generation module is configured to perform explicit thought chain reasoning and implicit thought chain reasoning on a sample question using the first language model to obtain a first sample reasoning process text and a first sample answer text for the sample question;
[0171] A second sample generation module is configured to perform explicit thought chain reasoning on the sample question using the second language model to obtain a second sample reasoning process text and a second sample answer text for the sample question;
[0172] A model training module is used to train the first large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample answer text, the second sample answer text and the true value of the answer to the sample question.
[0173] Optionally, the model training module is used to:
[0174] Determining a distillation loss between the first large language model and the second large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text;
[0175] determining a first inference loss of the first large language model based on the first sample reply text and the true value of the answer to the sample question;
[0176] The first large language model is trained based on the distillation loss and the first inference loss.
[0177] Optionally, the model training module is used to:
[0178] In a process in which the first large language model outputs the first sample reasoning process text and the first sample response text, determining a first hidden activation value of each hidden layer in the first large language model;
[0179] In a process in which the second large language model outputs the second sample reasoning process text and the second sample answer text, determining a second hidden activation value of each hidden layer in the second large language model;
[0180] The distillation loss is determined based on a difference between the aligned first hidden activation value and the second hidden activation value.
[0181] Optionally, the model training module is used to:
[0182] During the process of the first large language model outputting the first sample reasoning process text and the first sample reply text, determining a third hidden activation value in the first large language model, where the third hidden activation value is the last hidden activation value before the first large language model outputs the first sample reply text;
[0183] During the process of the second large language model outputting the second sample reasoning process text and the second sample reply text, determining a fourth hidden activation value in the second large language model, where the fourth hidden activation value is the last hidden activation value before the second large language model outputs the second sample reply text;
[0184] The distillation loss is determined based on a difference between the third hidden activation value and the fourth hidden activation value.
[0185] Optionally, the model training module is used to:
[0186] Determining a distillation loss between the first large language model and the second large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text;
[0187] determining a first inference loss of the first large language model based on the first sample reply text and the true value of the answer to the sample question;
[0188] determining a second inference loss of the second largest language model based on the second sample reply text and the true value of the answer to the sample question;
[0189] The first large language model and the second large language model are jointly trained based on the distillation loss, the first inference loss, and the second inference loss.
[0190] Optionally, the model training module is used to:
[0191] Based on the first loss weight, the second loss weight, and the third loss weight, fusing the distillation loss, the first reasoning loss, and the second reasoning loss to obtain a comprehensive reasoning loss;
[0192] Based on the comprehensive inference loss, the first language model and the second language model are jointly trained.
[0193] In summary, in the embodiment of the present application, the explicit thought chain and the implicit thought chain are combined and applied in the first language model, so that in the process of executing question reasoning, the first language model performs explicit thought chain reasoning on the input question, and the reasoning process text and the first reasoning vector of the input question can be obtained; the first language model performs implicit thought chain reasoning on the input question, and the second reasoning vector of the input question can be obtained. Finally, the answer text of the input question can be obtained based on the first reasoning vector and the second reasoning vector. The solution provided in the embodiment of the present application can ensure the accuracy of the first language model's question and answer of the input question based on the explicit thought chain, and at the same time improve the efficiency of the first language model's question and answer of the input question based on the implicit thought chain.
[0194] It should be noted that the apparatus provided in the above embodiments is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be distributed among different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.
[0195] Please refer to Figure 10 , which shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Specifically, the computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including a random access memory 1002 and a read-only memory 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 may also include a basic input / output system (I / O system) 1006 that helps transmit information between various components within the computer, and a large-capacity storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.
[0196] In some embodiments, the basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009, such as a mouse or keyboard, for user input. Both the display 1008 and the input device 1009 are connected to the central processing unit 1001 via an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may also include an input / output controller 1010 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.
[0197] The mass storage device 1007 is connected to the central processing unit 1001 via a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable media provide non-volatile storage for the computer device 1000. In other words, the mass storage device 1007 may include a computer-readable medium (not shown) such as a hard disk or drive.
[0198] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, tape cassette, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above-mentioned ones. The above-mentioned system memory 1004 and mass storage device 1007 can be collectively referred to as memory.
[0199] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 1001. The one or more programs contain instructions for implementing the above-mentioned method. The central processing unit 1001 executes the one or more programs to implement the question-answering method provided by the above-mentioned various method embodiments.
[0200] According to various embodiments of the present application, the computer device 1000 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1000 may be connected to the network 1011 via the network interface unit 1012 connected to the system bus 1005, or the network interface unit 1012 may be used to connect to other types of networks or remote computer systems (not shown).
[0201] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the question-answering method described in the above embodiment.
[0202] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs) or optical disks, etc. Among them, RAM may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).
[0203] An embodiment of the present application provides a computer program product, comprising at least one instruction stored in a computer-readable storage medium. A processor of a computer device reads the at least one instruction from the computer-readable storage medium and executes the at least one instruction, causing the computer device to perform the question-answering method described in the above embodiment.
[0204] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0205] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A question-answering method, characterized in that: The method comprises: Performing explicit thought chain reasoning on the input question through the first large language model to obtain a reasoning process text of the input question and a first reasoning vector, wherein the first reasoning vector is used to represent a result based on the explicit thought chain reasoning; Performing implicit thought chain reasoning on the input question using the first language model to obtain a second reasoning vector for the input question, where the second reasoning vector is used to represent a result based on the implicit thought chain reasoning; Based on the first inference vector and the second inference vector, generating a response text for the input question using the first large language model; outputting the reasoning process text and the answer text; Among them, the first language model is obtained based on the distillation learning of the second language model, and the second language model adopts the reasoning method based on the explicit thinking chain.
2. The method according to claim 1, characterized in that The step of performing explicit thought chain reasoning on the input question using the first large language model to obtain a reasoning process text of the input question and a first reasoning vector includes: In the process of performing explicit thought chain reasoning on the input question by using the first language model, each explicit reasoning step is discriminated by using the discriminant model to obtain a reasoning and discrimination result of each explicit reasoning step; Based on the reasoning judgment results of the respective explicit reasoning steps, a reasoning process text of the input question and the first reasoning vector are generated through the first large language model.
3. The method according to claim 2, characterized in that The reasoning process text includes reasoning text corresponding to n explicit reasoning steps, where n is a positive integer; In the process of performing explicit thought chain reasoning on the input question by using the first language model, discriminating each explicit reasoning step by using the discriminant model to obtain the reasoning and discrimination results of each explicit reasoning step includes: Performing an i-th step of explicit thought chain reasoning on the input question using the first language model to obtain an i-th step of reasoning text corresponding to the i-th explicit reasoning step, where i is a positive integer and is less than or equal to n; The i-th reasoning text is judged by a discriminant model to obtain an i-th reasoning judgment result of the i-th explicit reasoning step, and the i-th reasoning judgment result is used to indicate whether the i-th explicit reasoning step is accurate.
4. The method according to claim 3, characterized in that The generating, based on the reasoning judgment results of the respective explicit reasoning steps, the reasoning process text of the input question and the first reasoning vector by using the first large language model includes: When the i-th reasoning judgment result indicates that the i-th explicit reasoning step is accurate, performing the i+1-th explicit thought chain reasoning on the input question through the first large language model to obtain the i+1-th reasoning text of the input question; When the i-th reasoning judgment result indicates that the i-th explicit reasoning step is inaccurate, re-execute the i-th explicit thought chain reasoning on the input question using the first large language model to obtain the i-th reasoning text of the input question; When the n-1th reasoning judgment result indicates that the n-1th explicit reasoning step is accurate, the nth step of explicit thinking chain reasoning is performed on the input question through the first large language model to obtain the nth step reasoning text of the input question and the first reasoning vector.
5. The method according to any one of claims 1 to 4, characterized in that: Generating a response text to the input question using the first large language model based on the first inference vector and the second inference vector includes: Performing a linear transformation on the first inference vector and the second inference vector using the first large language model to obtain a query vector, a key vector, and a value vector; performing an attention calculation on the query vector, the key vector, and the value vector using the first language model to obtain an attention weight; The first language model performs weighted fusion on the attention weight and the value vector to generate the answer text of the input question.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Performing explicit thought chain reasoning and implicit thought chain reasoning on the sample question through the first language model to obtain a first sample reasoning process text and a first sample answer text of the sample question; Performing explicit thought chain reasoning on the sample question using the second language model to obtain a second sample reasoning process text and a second sample answer text for the sample question; The first large language model is trained based on the first sample reasoning process text, the second sample reasoning process text, the first sample answer text, the second sample answer text, and the true value of the answer to the sample question.
7. The method according to claim 6, characterized in that The training of the first large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample answer text, the second sample answer text, and the true value of the answer to the sample question includes: Determining a distillation loss between the first large language model and the second large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text; determining a first inference loss of the first large language model based on the first sample reply text and the true value of the answer to the sample question; The first large language model is trained based on the distillation loss and the first inference loss.
8. The method according to claim 7, characterized in that The determining, based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text, of a distillation loss between the first large language model and the second large language model includes: In a process in which the first large language model outputs the first sample reasoning process text and the first sample response text, determining a first hidden activation value of each hidden layer in the first large language model; In a process in which the second large language model outputs the second sample reasoning process text and the second sample answer text, determining a second hidden activation value of each hidden layer in the second large language model; The distillation loss is determined based on a difference between the aligned first hidden activation value and the second hidden activation value.
9. The method according to claim 7, characterized in that The determining, based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text, of a distillation loss between the first large language model and the second large language model includes: During the process of the first large language model outputting the first sample reasoning process text and the first sample reply text, determining a third hidden activation value in the first large language model, where the third hidden activation value is the last hidden activation value before the first large language model outputs the first sample reply text; During the process of the second large language model outputting the second sample reasoning process text and the second sample reply text, determining a fourth hidden activation value in the second large language model, where the fourth hidden activation value is the last hidden activation value before the second large language model outputs the second sample reply text; The distillation loss is determined based on a difference between the third hidden activation value and the fourth hidden activation value.
10. The method according to claim 6, characterized in that The training of the first large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample answer text, the second sample answer text, and the true value of the answer to the sample question includes: Determining a distillation loss between the first large language model and the second large language model based on the first sample reasoning process text, the second sample reasoning process text, the first sample reply text, and the second sample reply text; determining a first inference loss of the first large language model based on the first sample reply text and the true value of the answer to the sample question; determining a second inference loss of the second largest language model based on the second sample reply text and the true value of the answer to the sample question; The first large language model and the second large language model are jointly trained based on the distillation loss, the first inference loss, and the second inference loss.
11. The method according to claim 10, characterized in that The jointly training the first large language model and the second large language model based on the distillation loss, the first inference loss, and the second inference loss includes: Based on the first loss weight, the second loss weight, and the third loss weight, fusing the distillation loss, the first reasoning loss, and the second reasoning loss to obtain a comprehensive reasoning loss; Based on the comprehensive inference loss, the first language model and the second language model are jointly trained.
12. A question-answering device, characterized in that: The device comprises: a first reasoning module, configured to perform explicit chain of thought reasoning on an input question using a first large language model, to obtain a reasoning process text of the input question and a first reasoning vector, wherein the first reasoning vector is used to represent a result based on the explicit chain of thought reasoning; a second reasoning module, configured to perform implicit thought chain reasoning on the input question using the first large language model to obtain a second reasoning vector for the input question, wherein the second reasoning vector is used to represent a result based on the implicit thought chain reasoning; a text generation module, configured to generate a response text to the input question using the first large language model based on the first inference vector and the second inference vector; A text output module, configured to output the reasoning process text and the response text; Among them, the first language model is obtained based on the distillation learning of the second language model, and the second language model adopts the reasoning method based on the explicit thinking chain.
13. A computer device, characterized in that: The computer device includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the question-answering method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the question-answering method according to any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product includes at least one instruction, which is stored in a computer-readable storage medium; the processor of the computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device implements the question-answering method as described in any one of claims 1 to 11.
Citation Information
Cited By
Thinking chain optimization method and electronic equipment
CN121615788A