Model training method and device and electronic equipment
By introducing multiple thinking step instructions and training samples of sub-reply content in the automatic question-answer model, and using the base model to predict and update parameters, the problem of insufficient inference ability of the automatic question-answer model for complex problems is solved, and the accuracy of the reply content is improved.
Patent Information
- Application Number
- CN202510133132.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
The automatic question-and-answer model has weak inference ability for complex questions, resulting in low accuracy of reply content.
Through a model training method, a training sample set is obtained, wherein each sample includes a sample problem, multiple thinking step instructions, multiple sub-reply contents, and the first probability of each sub-reply content. The base model predicts the predicted probability of sub-reply content based on sample problems and thinking step instructions, determines the reply generation loss and parameter update loss, and updates the model parameters until the training end condition is reached.
The base model's ability to cope with complex problems has been improved, ensuring that the model can analyze and answer the answers in steps after training, and improving the accuracy of the reply content.
Smart Images

Figure CN120069072A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a model training method, apparatus, and electronic device. Background Art
[0002] With the rapid development of big data and computing power, neural network models have been widely used in various industries, including but not limited to image recognition, natural language processing, recommendation systems, question-and-answer systems, etc. In related technologies, the reasoning ability of automatic question-and-answer models for complex questions is weak, resulting in low accuracy of the reply content generated for complex questions. Summary of the Invention
[0003] In view of this, embodiments of this application propose a model training method, apparatus, and electronic device, which solve the problem that the reasoning ability of automatic question-and-answer models for complex questions in related technologies is weak.
[0004] Embodiments of this application are implemented by adopting the following technical solutions:
[0005] In a first aspect, an embodiment of this application provides a model training method, including: obtaining a training sample set, where each training sample in the training sample set includes a sample question, a plurality of thinking step instructions, a plurality of sub-reply contents, and a first probability of each sub-reply content, the plurality of thinking step instructions are used to represent a plurality of analysis steps for answering the sample question, the plurality of thinking step instructions correspond to the plurality of sub-reply contents one by one, and a sub-reply content is a reply content obtained by analyzing the sample question according to the analysis step represented by a thinking step instruction; the base model sequentially predicts the prediction probability of outputting the corresponding sub-reply content under each thinking step instruction of the sample question according to the sample question and the plurality of thinking step instructions of the sample question; determining a reply generation loss based on the prediction probabilities of the sub-reply contents respectively corresponding to the sample question under the plurality of thinking step instructions; determining a parameter update loss according to the first probability and the prediction probability of the sub-reply content respectively corresponding to the sample question under the plurality of thinking step instructions; and updating the parameters of the base model according to the reply generation loss and the parameter update loss until a training end condition is reached.
[0006] Second aspect, an embodiment of the present application provides a model training device, the device includes: a first acquisition module, configured to acquire a training sample set, each training sample in the training sample set includes a sample question, a plurality of thinking step instructions, a plurality of sub-response contents, and a first probability of each of the sub-response contents, the plurality of thinking step instructions are used to represent a plurality of analysis steps for answering the sample question, the plurality of thinking step instructions correspond to the plurality of sub-response contents one by one, and one sub-response content is a response content obtained by analyzing the sample question according to the analysis step represented by one thinking step instruction; a prediction module, configured to sequentially predict, by a base model, a prediction probability of outputting a corresponding sub-response content under each thinking step instruction of the sample question according to the sample question and the plurality of thinking step instructions of the sample question; a first loss calculation module, configured to determine a response generation loss based on the prediction probabilities of the sub-response contents respectively corresponding to the sample question under the plurality of thinking step instructions; a second loss calculation module, configured to determine a parameter update loss according to the first probability and the prediction probability of the sub-response contents respectively corresponding to the sample question under the plurality of thinking step instructions; a parameter adjustment module, configured to update the parameters of the base model according to the response generation loss and the parameter update loss until a training end condition is reached.
[0007] In some embodiments, the second loss calculation module includes a reward module, configured to predict, by a value network, an expected cumulative reward corresponding to the sub-response content of the sample question under the t-th thinking step instruction according to the sub-response content of the sample question under the t-th thinking step instruction; t is a positive integer starting from 1, t ≤ N, and N is the total number of thinking step instructions corresponding to the sample question; an advantage calculation module, configured to estimate an advantage function of the sub-response content of the sample question under the t-th thinking step instruction according to the expected cumulative reward corresponding to the sub-response content of the sample question under the t-th thinking step instruction and the prediction probability of the sub-response content of the sample question under the t-th thinking step instruction; a probability ratio calculation module, configured to determine a probability ratio of the sub-response content of the sample question under the t-th thinking step instruction according to the first probability and the prediction probability of the sub-response content of the sample question under the t-th thinking step instruction; an expectation calculation module, configured to calculate, according to the probability ratios and the advantage function of the sub-response contents of the sample question from the 1st thinking step instruction to the Nth thinking step instruction, an expectation of the base model for sequentially outputting the sub-response contents of the sample question from the 1st thinking step instruction to the Nth thinking step instruction as the parameter update loss.
[0008] In some embodiments, the desired computing module includes a first advantage return computing module for calculating the product of the probability ratio of the sub-response content of the sample problem at the t-th thinking step instruction and the corresponding advantage function to obtain the first advantage return of the sub-response content of the sample problem at the t-th thinking step instruction; a clipping module for clipping the probability ratio of the sub-response content of the sample problem at the t-th thinking step instruction according to a preset update amplitude interval to obtain the target probability ratio of the sub-response content of the sample problem at the t-th thinking step instruction; the target probability ratio is within the update amplitude interval; a second advantage return computing module for calculating the product of the advantage function of the sub-response content of the sample problem at the t-th thinking step instruction and the corresponding target probability ratio to obtain the second advantage return of the sub-response content of the sample problem at the t-th thinking step instruction; a selection module for taking the minimum value of the first advantage return and the second advantage return of the sub-response content of the sample problem at the t-th thinking step instruction as the target advantage return of the sub-response content of the sample problem at the t-th thinking step instruction; an accumulation module for calculating the expectation of the target advantage returns of the sub-response content of the sample problem from the first thinking step instruction to the N-th thinking step instruction to obtain the parameter update loss.
[0009] In some embodiments, when t is greater than 1, the base model predicts the prediction probability of the sample problem outputting the corresponding sub-response content at the t-th thinking step instruction according to the following process: the base model sequentially predicts the prediction probabilities of each character in the sub-response content output by the sample problem at the t-th thinking step instruction according to the sample problem, the sample problem from the first instruction response pair to the (t - 1)-th instruction response pair, and the t-th thinking step instruction of the sample problem; the (t - 1)-th instruction response pair includes the (t - 1)-th thinking step instruction of the sample problem and the sub-response content of the sample problem at the (t - 1)-th thinking step instruction; according to the prediction probabilities of each character in the sub-response content output by the sample problem at the t-th thinking step instruction, determine the prediction probability of the sample problem outputting the corresponding sub-response content at the t-th thinking step instruction; wherein, the base model sequentially predicts the prediction probabilities of each character in the sub-response content output by the sample problem at the first thinking step instruction according to the sample problem and the first thinking step instruction of the sample problem, and determines the prediction probability of the sample problem outputting the corresponding sub-response content at the first thinking step instruction.
[0010] In some embodiments, the model training device further includes a data preparation module for obtaining an initial data set, where the initial data set includes a plurality of reference questions and the reference reply content corresponding to each reference question; an instruction generation module for generating, by a chain of thought generation model, a plurality of thought step instructions for the reference questions; the plurality of thought step instructions for the reference questions being used to represent a plurality of analysis steps for answering the reference questions; a reference reply module for generating a reply by a reference base model according to the reference question and the plurality of thought step instructions for the reference question, to obtain candidate sub-reply content for the reference question under the plurality of thought step instructions; an addition module for, if the semantic similarity between the candidate sub-reply content of the reference question under the corresponding last thought step instruction and the reference reply content corresponding to the reference question exceeds a similarity threshold, taking the reference question, the plurality of thought step instructions for the reference question, and the candidate sub-reply content respectively corresponding to the sample under the plurality of thought step instructions as a training sample and adding it to the training sample set.
[0011] In some embodiments, the model training device further includes an execution module for, if the semantic similarity between the candidate sub-reply content of the reference question under the corresponding last thought step instruction and the reference reply content corresponding to the reference question does not exceed the similarity threshold, guiding the reference base model to regenerate a reply according to the reference question and the plurality of thought step instructions for the reference question.
[0012] In some embodiments, the model training device further includes: a second acquisition module for acquiring a target question; a chain of thought generation module for generating, by a chain of thought generation model based on the target question, a plurality of thought step instructions for the target question; a reply module for generating a reply by a base model according to the target question and the plurality of thought step instructions for the target question, to obtain target sub-reply content for the target question under the plurality of thought step instructions; an output module for taking the target sub-reply content of the target question under the corresponding last thought step instruction as the target reply content of the target question.
[0013] In some embodiments, the response module is specifically configured to cause the base model to output the target sub-response content of the target question under the first thinking step instruction according to the target question and the first thinking step instruction for the target question; cause the base model to predict the target sub-response content of the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction of the target question; wherein, the (i - 1)-th instruction response pair includes the (i - 1)-th thinking step instruction for the target question and the target sub-response content of the target question under the (i - 1)-th thinking step instruction; wherein, i is a positive integer greater than 1 and less than M, and M is the total number of thinking step instructions corresponding to the target question; determine the accuracy score of the target sub-response content of the target question under the i-th thinking step instruction; if the accuracy score of the target sub-response content of the target question under the i-th thinking step instruction is greater than the score threshold, increment i by 1, and return to execute causing the base model to predict the target sub-response content of the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction of the target question.
[0014] In some embodiments, the response module is further configured to, if the accuracy score of the target sub-response content of the target question under the i-th thinking step instruction is not greater than the score threshold, increment the number of response times corresponding to the i-th thinking step instruction by 1; if the number of response times corresponding to the i-th thinking step instruction is less than the response times threshold, guide the base model to regenerate the target sub-response content output by the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction of the target question; if the number of response times corresponding to the i-th thinking step instruction is not less than the response times threshold, increment i by 1, and return to execute causing the base model to predict the target sub-response content of the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction of the target question.
[0015] In some embodiments, the thought chain generation module includes an expansion module for generating thought step instructions according to the target problem by the thought chain model to obtain a tree-shaped thought chain of the target problem. A tree-shaped thought chain of a target problem includes multiple thought step instruction branches, and each thought step instruction branch contains multiple thought step instructions of the target problem from the front to the back; a screening module for calculating the path score of each thought step instruction branch and determining a target thought step instruction branch from the multiple thought step instruction branches included in the tree-shaped thought chain of the target problem based on the path scores of the thought step instruction branches of each target problem; a determination module for using the multiple thought step instructions included in the target thought step instruction branch as the multiple thought step instructions of the target problem.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor; a memory, on which computer instructions are stored, and when the computer instructions are executed by the processor, the above method is implemented.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a processor, the above method is implemented.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product including computer instructions, and when the computer instructions are executed by a processor, the above method is implemented.
[0019] In this application, through sample questions, multiple step-by-step thinking instruction for the sample questions, sub-answer contents respectively corresponding to the sample questions under multiple thinking instructions, and the first probability of each sub-answer content; the base model sequentially predicts the prediction probability of outputting the corresponding sub-answer content under each thinking instruction of the sample question according to the sample question and the multiple thinking instructions of the sample question; based on the prediction probability of the sub-answer content respectively corresponding to the sample question under multiple thinking instructions, the reply generation loss is determined; according to the first probability and the prediction probability of the sub-answer content respectively corresponding to the sample question under multiple thinking instructions, the parameter update loss is determined; the reply generation loss can supervise the base model to output accurate sub-answer contents under each thinking instruction of the sample question; the parameter update loss is used to control the update amplitude of the parameters of the base model, making the learning process of the base model more stable. Since the base model is pre-trained through a large amount of data and already has good semantic understanding and question-answering capabilities, introducing the parameter update loss can avoid the situation where the performance of the base model drops sharply during the training process. In addition, in this application, using the sample question, multiple step-by-step thinking instructions of the sample question, sub-answer contents respectively corresponding to the sample question under multiple thinking instructions, and the first probability of each sub-answer content to train the base model can enable the base model to learn the ability to split problems step by step and answer them gradually. In this way, the ability of the base model to handle complex problems can be effectively improved, and it can be ensured that after training, the base model can also analyze and answer complex problems step by step, improving the accuracy of the output reply content.
[0020] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained without creative efforts based on these drawings.
[0022] Figure 1 Fig. shows a schematic diagram of an application scenario related to an embodiment of the present application.
[0023] Figure 2 Fig. shows a schematic flowchart of a model training method provided by an embodiment of the present application.
[0024] Figure 3 Fig. shows the one provided by an embodiment of the present application Figure 2 Schematic flowchart of step S120.
[0025] Figure 4 shows the process schematic diagram of step S140 provided in an embodiment of the present application. Figure 2 in
[0026] Figure 5 shows the process schematic diagram of step S340 provided in an embodiment of the present application. Figure 4 in
[0027] Figure 6 shows another process schematic diagram of the model training method provided in an embodiment of the present application.
[0028] Figure 7 shows the process schematic diagram of the model training method provided in another embodiment of the present application.
[0029] Figure 8 shows the process schematic diagram of step S630 provided in an embodiment of the present application. Figure 7 in
[0030] Figure 9 shows another process schematic diagram of the response generation method provided in an embodiment of the present application.
[0031] Figure 10 shows the process schematic diagram of step S620 provided in an embodiment of the present application. Figure 7 in
[0032] Figure 11 shows the schematic diagram of the model training device provided in an embodiment of the present application.
[0033] Figure 12 shows the schematic diagram of the electronic device provided in an embodiment of the present application. Detailed Embodiments
[0034] The following describes in detail the embodiments of the present application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.
[0035] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0036] In the following description, the terms "first", "second", etc. involved are only used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first" and "second" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In the following description, when referring to "some embodiments or some implementation manners", it describes a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0037] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0038] For the convenience of understanding the present application, some terms will be explained first below.
[0039] Chain of Thought (CoT): A prompting technique that guides a language model to generate intermediate reasoning steps; by generating the reasoning process, it improves the model's ability to solve complex problems and interpretability.
[0040] Implicit Chain of Thought: A thought step instruction that does not directly show intermediate reasoning steps, but has an implicit intermediate reasoning step process inside the model; compared with the explicit chain of thought, it avoids the process of manually annotating intermediate steps.
[0041] Base Model: A general language model that has been pre-trained on a large scale and can be further fine-tuned for specific tasks. It has extensive knowledge and language understanding capabilities, but may need to be optimized for specific tasks.
[0042] With the rapid development of big data and computing power, neural network models have been widely used in various industries, including but not limited to image recognition, natural language processing, recommendation systems, question-and-answer systems, etc. In related technologies, the reasoning ability of automatic question-and-answer models for complex questions is weak, resulting in low accuracy of the reply content generated for complex questions.
[0043] To solve the above problems, the present application provides a model training method, a reply generation method, a device, and an electronic device.
[0044] Please refer to Figure 1 , Figure 1The application scenario diagram involved in the embodiments of the present application is shown. This application scenario includes a terminal 10 and a server 20, where the terminal 10 and the server 20 are connected via a wired or wireless network.
[0045] The terminal 10 can install application programs. A base model is deployed in the server 20. Users can input a target question through the application programs installed on the terminal 10. Subsequently, the application programs call the base model deployed in the server 20. The base model can automatically generate the target reply content of the target question according to the method provided in the present application and send the target reply content to the terminal 10.
[0046] Among them, the base model deployed in the server 20 can be trained according to the method provided in the present application. Specifically, a training sample set is obtained. Each training sample in the training sample set includes a sample question, multiple step-by-step thinking instruction of the sample question from the front to the back, sub-reply content respectively corresponding to the sample question under multiple step-by-step thinking instructions, and the first probability of each sub-reply content; the base model predicts the prediction probability of outputting the corresponding sub-reply content under each step-by-step thinking instruction of the sample question according to the sample question and the multiple step-by-step thinking instructions of the sample question in sequence; based on the prediction probability of the sub-reply content respectively corresponding to the sample question under multiple step-by-step thinking instructions, the reply generation loss is determined; according to the first probability and the prediction probability of the sub-reply content respectively corresponding to the sample question under multiple step-by-step thinking instructions, the parameter update loss is determined; according to the reply generation loss and the parameter update loss, the parameters of the base model are updated until the training end condition is reached.
[0047] In some other embodiments, the base model can also be deployed on the terminal 10, and the terminal 10 trains the base model according to the method provided in the present application.
[0048] Among them, the terminal 10 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a wearable device, a smart watch, a virtual reality device, a vehicle-mounted terminal, etc., but is not limited thereto.
[0049] The server 20 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0050] Please refer to Figure 2 , Figure 2 The flow diagram of the model training method provided by the embodiments of the present application is shown. This method can be executed by an electronic device, and the electronic device can be a server, a terminal, etc., which is not specifically limited herein. AsFigure 2 As shown, the method includes steps S110 - S150:
[0051] S110. Obtain a training sample set, where each training sample in the training sample set includes a sample question, multiple thinking step instructions, multiple sub - reply contents, and a first probability of each of the sub - reply contents. The multiple thinking step instructions are used to represent multiple analysis steps for answering the sample question. The multiple thinking step instructions and the multiple sub - reply contents are in one - to - one correspondence. One sub - reply content is the reply content obtained by analyzing the sample question in the analysis step represented by one thinking step instruction.
[0052] Among them, the sample question can be pure text content or multi - modal content, such as the combination of image content and text content, the combination of audio content and text content, the combination of audio - video content, etc., and no specific limitation is made here.
[0053] The multiple thinking step instructions of the sample question refer to text instructions representing multiple analysis steps for answering the sample question. One thinking step instruction of the sample question represents one analysis step. The sub - reply content of the sample question under one thinking step instruction refers to the answer result obtained by analyzing the sample question according to the thinking step instruction. It can be understood that the multiple thinking step instructions of the sample question represent multiple analysis steps for answering the sample question. Therefore, the multiple thinking step instructions of the sample question have a sequential order. Furthermore, the sub - reply content of the sample question under one thinking step instruction (assumed to be thinking step instruction A) is determined by combining the sample question and the sub - reply content of the sample question under the thinking step instructions before thinking step instruction A and answering according to thinking step instruction A. In some embodiments, the multiple thinking step instructions of the sample question can be manually marked by testers for the sample question; or they can be generated by a thought - chain model according to the sample question to obtain multiple thinking step instructions from first to last.
[0054] In some embodiments, the sub - reply contents respectively corresponding to the sample question under multiple thinking step instructions can be manually marked by testers. On this basis, the sample question, the multiple thinking step instructions of the sample question from first to last, and the sub - reply contents respectively corresponding to the sample question under multiple thinking step instructions can be input into a language model for prediction to obtain the first probability of each sub - reply content.
[0055] In some other embodiments, it is also possible to input the sample problem and multiple step-by-step thinking instruction for the sample problem into the reference base model for prediction, so as to obtain the sub-response content corresponding to the sample problem under each of the multiple step-by-step thinking instructions and the first probability of each sub-response content. The reference base model can be the base model to be trained in this application, or other large language models with text generation functions.
[0056] On this basis, since the sub-response content of the sample problem generated by the reference base model may contain incorrect response content, it is also possible to judge the accuracy of the generated sub-response content. For example, it is possible to judge the accuracy of the sub-response content under each step-by-step thinking instruction output by the reference base model. If the accuracy of the sub-response content under a certain step-by-step thinking instruction is low, the sub-response content under this step-by-step thinking instruction can be regenerated. It is also possible to only judge the accuracy of the sub-response content corresponding to the sample problem under the last step-by-step thinking instruction. If the sub-response content corresponding to the sample problem under the last step-by-step thinking instruction is accurate, it means that the sub-response content corresponding to other step-by-step thinking instructions is also probably accurate. If the sub-response content corresponding to the sample problem under the last step-by-step thinking instruction is inaccurate, it means that the sub-response content corresponding to other step-by-step thinking instructions is also probably inaccurate. At this time, this training sample can be discarded, or the sub-response content corresponding to the sample problem under each of the multiple step-by-step thinking instructions can be regenerated until the accuracy of the sub-response content corresponding to the sample problem under the last step-by-step thinking instruction is judged to be accurate.
[0057] Among them, the accuracy of the sub-response content of the sample problem under a certain step-by-step thinking instruction can be obtained by calling an external discriminant model and inputting the sample problem, the multiple step-by-step thinking instructions of the sample problem from the first to the last, and the sub-response content into the discriminant model, and the discriminant model outputs the accuracy of the sub-response content. It can also be directly determined according to the prediction probabilities of each character in the sub-response content. For example, the prediction probabilities of each character in the sub-response content are weighted and summed, and the summation result is used as the accuracy of the sub-response content.
[0058] S120. The base model sequentially predicts the prediction probabilities of the corresponding sub-response content output under each step-by-step thinking instruction of the sample problem according to the sample problem and the multiple step-by-step thinking instructions of the sample problem.
[0059] Among them, the base model can be a neural network model for reply generation, which can adopt a transformer architecture. Specifically, the base model can include an input embedding layer, a position encoding layer, a hidden layer, and an output layer.
[0060] The process of the base model for reply generation can include the following steps 1 to 4:
[0061] Step 1: Each character in the input text is embedded by the input embedding layer to obtain the embedding vectors of the characters in the input text.
[0062] Exemplarily, assume that the vocabulary includes v characters, and the dimension of the embedding vector of each character in the vocabulary is d, that is, the embedding vocabulary can be represented as a vocabulary matrix E ∈ R^(v×d); for the input sequence x = (x_1, …, x_T), performing embedding processing on it can obtain the embedding vector matrix h_0 = [Ex_1, …, Ex_T], where Ex_i is the embedding vector retrieved according to the character x_i in the input text from the vocabulary matrix E.
[0063] Step 2: Each character in the input sequence is position-encoded by the position encoding layer, and the position encoding of each character is added to the corresponding embedding vector, so that the position information of each character in the input sequence can be captured.
[0064] Exemplarily, the sine and cosine functions can be used to perform position encoding on each character in the input sequence, that is, for the positions of even dimensions, its position encoding is: PE(pos, 2i) = sin(pos / 10000^(2i / d)), and for the positions of odd dimensions, its position encoding is: PE(pos, 2i + 1) = cos(pos / 10000^(2i / d)); where pos represents the position of the character in the input sequence, i is the i-th dimension of the embedding vector, and d is the total dimension of the embedding vector of the character.
[0065] After that, the position encoding of each character is added to the corresponding embedding vector to obtain a vector with position information; that is, for the embedding vector matrix h_0 = [Ex_1, …, Ex_T] of the input sequence x = (x_1, …, x_T), after adding the position encoding, a new input vector matrix h_1 = [Ex_1 + PE(1), …, Ex_T + PE(T)] can be obtained, where PE(i) is the position encoding of the character x_i.
[0066] Step 3: The hidden layer extracts features based on the input vectors (including embedding vectors and position encoding) of the input sequence and outputs the hidden layer features.
[0067] Exemplarily, the hidden layer may include multiple layers of Transformer blocks, and each layer of Transformer blocks includes a multi-head self-attention layer and a feed-forward neural network. Suppose there are L layers of Transformer blocks, and the calculation process of the Transformer block in the L-th layer is h_L = TransformerBlock(h_(L-1)); where h_(L-1) is the output of the Transformer block in the (L-1)-th layer, and h_L is the output of the Transformer block in the L-th layer, that is, the output of each layer of Transformer blocks takes the output of the previous layer of Transformer blocks as the input.
[0068] Among them, each layer of Transformer blocks includes the following operations:
[0069] Calculate multi-head attention: The input vector h after position encoding is respectively mapped to the query, key, and value spaces through three different linear transformations (weight matrices are W_i^ Q , W_i^ K , W_i^ V ), that is, Q_i, K_i, V_i = W_i^ Q h, W_i^ K h, W_i^ V h; then calculate the attention scores and weighted sum to obtain the output of each attention head, that is, the attention scores Among them, d_k is the dimension of the key vector, which is used to scale the attention scores to prevent gradient vanishing or explosion; finally, the outputs of all heads are concatenated and passed through a linear transformation (weight matrix is W^ O ) to obtain the final output of the multi-head self-attention layer, that is, multi-head attention MultiHead(h) = Concat(head_1,..., head_h)W^ O .
[0070] After that, the feed-forward neural network takes the output of the multi-head self-attention layer as the input, and after being processed by a non-linear activation function and a linear transformation, outputs an intermediate vector. Taking the non-linear activation function as the ReLU (Rectified Linear Unit) function as an example, the output intermediate vector FFN(x) = max(0, xW_1 + b_1)W_2 + b_2; where x represents the input of the feed-forward neural network, that is, the output of the multi-head self-attention layer, and W_1 and W_2 are both weight matrices of the feed-forward neural network; b_1 and b_2 are both bias vectors of the feed-forward neural network.
[0071] During the above processing, the output of the multi-head self-attention layer can be output to the feed-forward neural network after layer normalization and residual connection processing; the output of the feed-forward neural network can also be output to the next Transformer block after layer normalization and residual connection processing, thereby improving the training effect of the model.
[0072] Specifically, layer normalization and residual connection processing are performed on the multi-head self-attention MultiHead(h), that is, h’ = LayerNorm(h + MultiHead(h)); then, layer normalization and residual connection processing are performed on the output of the feed-forward neural network, that is, h_out = LayerNorm(h’ + FFN(h’)).
[0073] Step 4: The output layer maps the output of the last Transformer block to the vocabulary space to obtain the probability of each character in the vocabulary space as the output character.
[0074] Specifically, the output of the last Transformer block is mapped to a vector logits of the vocabulary size, that is, logits = h_LW_out + b_out; where h_L is the output of the last Transformer block, W_out is a weight matrix, b_out is a bias vector, and logits is a T×v matrix, where the t-th row represents the unnormalized score of each character in the vocabulary corresponding to the t-th position in the sequence.
[0075] The vector logits is converted into a probability distribution using the softmax function, that is, p(x_t|x<t) = softmax(logits_t), where logits_t is the mapping vector of the t-th character, x_t represents the character at the t-th position of the output, x<t represents all characters before the t-th position of the output, and p(x_t|x<t) represents the probability that the t-th position is the character x_t in the vocabulary given x<t.
[0076] Among them, the predicted probability of the sub-response content corresponding to the sample problem under multiple thinking step instructions can be determined by the predicted probability of each character in the sub-response content; for example, the predicted probability of a sub-response content can be the sum of the predicted probabilities of all characters in the sub-response content, or the mean of the predicted probabilities of all characters in the sub-response content. The predicted probability of the sub-response content of the sample problem under one thinking step instruction is used to characterize the confidence of the sub-response content of the sample problem under one thinking step instruction. The higher the predicted probability, the higher the credibility of the sub-response content of the sample problem under one thinking step instruction.
[0077] As described in the above process, during the process of the base model making a prediction response, it will predict the probability of each character in the vocabulary space as the output character and sample according to the probability. Therefore, in the case where the sample question, multiple thinking step instructions of the sample question, and the sub-response content corresponding to the sample question under multiple thinking step instructions have been determined, the base model can determine the prediction probability of each character in the sub-response content as the output character at the corresponding position, and then determine the prediction probability of the sub-response content according to the prediction probability corresponding to each character in the sub-response content.
[0078] For example, for the first thinking step instruction of the sample question, the base model can sequentially predict the prediction probability of each character in the sub-response content corresponding to the output of the sample question under the first thinking step instruction according to the sample question and the first thinking step instruction of the sample question, and determine the prediction probability of the sub-response content corresponding to the output of the sample question under the first thinking step instruction.
[0079] For example, in the training sample, if the first thinking step instruction of the sample question Q1 is X1, and the sub-response content corresponding to the first thinking step instruction of the sample question Q1 is Y = {y1, y2, y3, y4,..., yn}, where ya refers to the a-th character in the sub-response content corresponding to the first thinking step instruction of the sample question Q1, a is a positive integer, and n is the total number of characters in the sub-response content corresponding to the first thinking step instruction of the sample question Q1, 1 ≤ a ≤ n.
[0080] The sample question Q1 and the first thinking step instruction X1 of the sample question Q1 can be input into the base model. The base model can jointly semantically encode the sample question Q1 and the first thinking step instruction X1 of the sample question Q1 to obtain semantic encoding features. Then, the semantic encoding features are decoded character by character to determine the probability that the a-th character is ya. It should be noted that when a > 1, the base model combines the semantic encoding features and the embedding vectors of the first a - 1 characters in the sub-response content Y corresponding to the first thinking step instruction of the sample question Q1 for decoding and outputs the probability that the a-th character is ya. According to the above process, the base model can sequentially output the probability that the first character is y1, the probability that the second character is y2, the probability that the third character is y3.... to the probability that the n-th character is yn.
[0081] In some embodiments, the prediction probabilities of each character in the sub-response content corresponding to the output of the sample question under the first thinking step instruction can be averaged, and the obtained average prediction probability can be used as the prediction probability of the sub-response content corresponding to the output of the sample question under the first thinking step instruction.
[0082] Please refer to Figure 3 ,Figure 3 The present application provides an Figure 2 Schematic flowchart of step S120 in []. When the serial number t of the thinking step instruction is greater than 1, the base model predicts the prediction probability of the corresponding sub-response content output by the sample question under the t-th thinking step instruction according to the following process. Step S130 may include steps S210-S220:
[0083] S210. The base model sequentially predicts the prediction probability of each character in the corresponding sub-response content output by the sample question under the t-th thinking step instruction according to the sample question, the sample question from the 1st instruction response pair to the (t-1)-th instruction response pair, and the t-th thinking step instruction of the sample question; the (t-1)-th instruction response pair includes the t-1-th thinking step instruction of the sample question and the sub-response content of the sample question under the t-1-th thinking step instruction.
[0084] Among them, the sample question from the 1st instruction response pair to the (t-1)-th instruction response pair is used to provide the basic knowledge for the base model to answer the sample question under the t-th thinking step instruction. That is, the base model is equivalent to providing the answer content of the sample question under the t-th thinking step instruction on the basis that the answer content of the sample question under the previous t-1 thinking step instructions is the corresponding sub-response content.
[0085] Similarly, the base model can jointly semantically encode the sample question, the sample question from the 1st instruction response pair to the (t-1)-th instruction response pair, and the t-th thinking step instruction of the sample question to obtain a semantic encoding feature. Then, based on the semantic encoding feature, character-by-character decoding is performed to obtain the probability that the b-th character output is the b-th character in the sub-response content of the sample question under the t-th thinking step instruction, where b is a positive integer. The process of character-by-character decoding is similar to the above process and will not be elaborated here.
[0086] S220. Determine the prediction probability of the corresponding sub-response content output by the sample question under the t-th thinking step instruction according to the prediction probability of each character in the corresponding sub-response content output by the sample question under the t-th thinking step instruction.
[0087] In some embodiments, it may be to perform weighted summation (for example, calculating the mean) on the prediction probability of each character in the corresponding sub-response content output by the sample question under the t-th thinking step instruction to obtain the prediction probability of the corresponding sub-response content output by the sample question under the t-th thinking step instruction; it may also be to perform weighted summation on the prediction probability of each character in the sub-response content output by the sample question under the 1st thinking step instruction to the t-th thinking step instruction respectively to obtain the prediction probability of the corresponding sub-response content output by the sample question under the t-th thinking step instruction.
[0088] In the above embodiments, except that the prediction probability of the sub-response content under the first thinking step instruction is determined according to the sample question and the first thinking step instruction of the sample question; for the prediction probability of the sub-response content under the t-th thinking step instruction, the base model determines it according to the sample question, the sample question from the first instruction response pair to the t-1-th instruction response pair, and the t-th thinking step instruction of the sample question, which can guide the base model to learn the ability to reason step by step for the question and then generate a response, and improve the reasoning ability of the base model.
[0089] S130. Determine the response generation loss based on the prediction probabilities of the sub-response contents respectively corresponding to the sample question under multiple thinking step instructions.
[0090] In some embodiments, the response generation loss can be calculated by a loss function, such as a log-likelihood function, cross-entropy loss, etc.
[0091] Taking the loss function as the log-likelihood function as an example, the response generation loss L_LM = -∑t log p(x_t|x<t); where ∑ is the summation symbol, t represents the serial number of the thinking step instruction, x_t represents the sub-response content of the sample question under the t-th thinking step instruction, x<t represents the input for generating the sub-response content under the t-th thinking step instruction, and p(x_t|x<t) represents the probability of outputting x_t when the input is x<t. Among them, when t is 1, x<t represents the sample question and the first thinking step instruction, and when t is greater than 1, x<t represents the sample question, the sample question from the first instruction response pair to the t-1-th instruction response pair, and the t-th thinking step instruction of the sample question, where the t-1-th instruction response pair includes the t-1-th thinking step instruction for the sample question and the sub-response content of the sample question under the t-1-th thinking step instruction.
[0092] S140. Determine the parameter update loss according to the first probability and the prediction probability of the sub-response contents respectively corresponding to the sample question under multiple thinking step instructions.
[0093] Among them, the parameter update loss is used to limit the update amplitude of the base model during training, make the learning process of the base model more stable, and avoid the situation where the performance of the base model drops sharply.
[0094] In some embodiments, please refer to Figure 4 , Figure 4 which gives the flow schematic diagram of step S140 provided by the embodiment of the present application. Step S140 may include steps S310 - S340: Figure 2
[0095] S310. The value network predicts the expected cumulative reward corresponding to the sub - reply content of the sample question under the t - th thinking step instruction according to the sub - reply content of the sample question under the t - th thinking step instruction. t is a positive integer starting from 1, t ≤ N, and N is the total number of thinking step instructions corresponding to the sample question.
[0096] Among them, the value network can be a pre - trained reward model, which is used to predict the expected cumulative reward corresponding to the sub - reply content of the sample question under the t - th thinking step instruction according to the sub - reply content of the sample question under the t - th thinking step instruction. The expected cumulative reward corresponding to the sub - reply content of the sample question under the t - th thinking step instruction refers to the sum of the rewards expected to be obtained from generating the sub - reply content of the sample question under the t - th thinking step instruction to generating the sub - reply content of the sample question under the N - th thinking step instruction.
[0097] S320. Estimate the advantage function of the sub - reply content of the sample question under the t - th thinking step instruction according to the expected cumulative reward corresponding to the sub - reply content of the sample question under the t - th thinking step instruction and the predicted probability of the sub - reply content of the sample question under the t - th thinking step instruction.
[0098] Among them, the advantage function of the sub - reply content of the sample question under the t - th thinking step instruction is used to evaluate the advantage of generating the corresponding sub - reply content compared with other reply contents under the t - th thinking step instruction for the sample question. Further, according to the predicted probability of the sub - reply content of the sample question under the t - th thinking step instruction, the probability of generating other sub - reply contents under the t - th thinking step instruction can be determined. If the expected cumulative reward corresponding to the sub - reply content of the sample question under the t - th thinking step instruction is greater than the average cumulative reward of generating other sub - reply contents of the sample question under the t - th thinking step instruction, it is considered that generating this sub - reply content is advantageous under the t - th thinking step instruction. Obviously, since the base model always outputs the optimal solution as the reply content during application, the higher the advantage function corresponding to a sub - reply content, the easier it is for the base model to output this sub - reply content, and this sub - reply content is also better.
[0099] S330. Determine the probability ratio of the sub - reply content of the sample question under the t - th thinking step instruction according to the first probability and the predicted probability of the sub - reply content of the sample question under the t - th thinking step instruction.
[0100] Among them, the first probability of the sub - reply content of the sample question under the t - th thinking step instruction can be determined by the reference base model.
[0101] The closer the probability ratio is to 1, it represents that the difference between the probability distributions of the sub - reply contents output by the currently trained base model and the reference base model for the sample problem under the t - th thinking step instruction is small, that is, the difference between the currently trained base model and the reference base model is small; the farther the probability ratio is from 1, it represents that the difference between the probability distributions of the sub - reply contents output by the currently trained base model and the reference base model for the sample problem under the t - th thinking step instruction is large, that is, the difference between the currently trained base model and the reference base model is large.
[0102] S340. Calculate the expectation of the sub - reply contents output by the base model for the sample problem in sequence from the 1st thinking step instruction to the Nth thinking step instruction according to the probability ratio and the advantage function of the sub - reply contents for the sample problem under the 1st thinking step instruction to the Nth thinking step instruction, and use it as the parameter update loss.
[0103] In the above - mentioned embodiment, the difference between the currently trained base model and the reference base model can be determined through the probability ratio; the quality of the currently trained base model can be determined through the advantage function. The parameter update loss obtained based on the probability ratio and the advantage function can consider both the quality of the currently trained base model and the update amplitude of the currently trained base model. When adjusting the base model using the parameter update loss, it can effectively avoid the training instability caused by too large an update amplitude and ensure the stability of the base model.
[0104] In some embodiments, please refer to Figure 5 , Figure 5 shows the process schematic diagram of step S340 provided by the embodiment of the present application. Step S340 may include steps S410 - S450: Figure 4 In
[0105] S410. Calculate the product of the probability ratio of the sub - reply content for the sample problem under the t - th thinking step instruction and the corresponding advantage function to obtain the first advantage return of the sub - reply content for the sample problem under the t - th thinking step instruction.
[0106] S420. According to the preset update amplitude interval, perform a clipping operation on the probability ratio of the sub - reply content for the sample problem under the t - th thinking step instruction to obtain the target probability ratio of the sub - reply content for the sample problem under the t - th thinking step instruction; the target probability ratio is within the update amplitude interval.
[0107] Specifically, when the probability ratio is within the preset update range interval, no clipping operation is required for the probability ratio, that is, the target probability ratio is the probability ratio; when the probability ratio is not within the preset update range interval, a clipping operation is required for the probability ratio. If the probability ratio is greater than the maximum value of the preset update range interval, the maximum value of the preset update range interval is used as the target probability ratio; if the probability ratio is less than the minimum value of the preset update range interval, the minimum value of the preset update range interval is used as the target probability ratio.
[0108] Obviously, by performing a clipping operation on the probability ratio, the range of the probability ratio is effectively restricted, which also restricts the update range of the current training base model.
[0109] S430. Calculate the product of the advantage function of the sub-response content of the t-th thinking step instruction for the sample problem and the corresponding target probability ratio to obtain the second advantage return of the sub-response content of the t-th thinking step instruction for the sample problem.
[0110] S440. Take the minimum value of the first advantage return and the second advantage return of the sub-response content of the t-th thinking step instruction for the sample problem as the target advantage return of the sub-response content of the t-th thinking step instruction for the sample problem.
[0111] S450. Calculate the expectation of the target advantage returns of the sub-response content of the sample problem from the first thinking step instruction to the N-th thinking step instruction to obtain the parameter update loss.
[0112] Specifically, the parameter update loss L^CLIP(θ) can be calculated by the following formula:
[0113] L^CLIP(θ) = E_t[min(r_t(θ)A_t, clip(r_t(θ), 1 - ε, 1 + ε)A_t)].
[0114] Where θ is the model parameter, E represents the calculation of expectation, t represents the serial number of the thinking step instruction for the sample problem, r_t(θ) represents the probability ratio of the sub-response content of the t-th thinking step instruction for the sample problem, A_t is the advantage function of the sub-response content of the t-th thinking step instruction for the sample problem, clip represents the clipping function, clip(r_t(θ), 1 - ε, 1 + ε) is also the target probability ratio, which means that when r_t(θ) is less than 1 - ε or greater than 1 + ε, it is clipped to be compressed within the range of [1 - ε, 1 + ε]. [1 - ε, 1 + ε] is the update range interval in the above text, and ε is a small constant used to limit the update range of the parameters of the base model.
[0115] Among them, \(r_t(\theta)=\frac{\pi_{\theta}(a_t|s_t)}{\pi_{\theta_{old}}(a_t|s_t)}\), where \(\pi_{\theta}(a_t|s_t)\) represents the probability of the currently trained base model performing action \(a_t\) in the case of input state \(s_t\), that is, the prediction probability of the sub - reply content of the currently trained base model for the sample problem under the \(t\) - th thinking step instruction; \(\pi_{\theta_{old}}(a_t|s_t)\) refers to the probability of the reference base model performing action \(a_t\) in the case of input state \(s_t\), that is, the first probability of the sub - reply content of the reference base model for the sample problem under the \(t\) - th thinking step instruction.
[0116] S150. Update the parameters of the base model according to the reply generation loss and the parameter update loss until the training end condition is reached.
[0117] Specifically, it can be to perform weighted summation on the reply generation loss and the parameter update loss to obtain the target loss, that is, the target loss \(L_{total}=\alpha L_{LM}+(1 - \alpha)L^{CLIP}\), where \(\alpha\) is a trade - off parameter used to adjust the relative importance of the two reply generation losses and the parameter update loss. Then, update the parameters of the base model according to the target loss until the training end condition is reached.
[0118] The training end condition can be that the target loss is less than the loss threshold, or that the number of training iterations is greater than the number threshold.
[0119] The model training method provided by this application includes a sample question, multiple step-by-step thinking instruction for the sample question from the first to the last, sub-answer content respectively corresponding to the sample question under multiple thinking instructions, and the first probability of each sub-answer content. The base model sequentially predicts the prediction probability of outputting the corresponding sub-answer content under each thinking instruction of the sample question according to the sample question and multiple thinking instructions of the sample question. Based on the prediction probability of the sub-answer content respectively corresponding to the sample question under multiple thinking instructions, the reply generation loss is determined. According to the first probability and the prediction probability of the sub-answer content respectively corresponding to the sample question under multiple thinking instructions, the parameter update loss is determined. The reply generation loss can supervise the base model to output accurate sub-answer content under each thinking instruction of the sample question. The parameter update loss is used to control the update amplitude of the parameters of the base model, making the learning process of the base model more stable. Since the base model is pre-trained with a large amount of data and already has good semantic understanding and question-answering capabilities, introducing the parameter update loss can avoid the situation where the performance of the base model drops sharply during training. In addition, in this application, using the sample question, multiple step-by-step thinking instructions for the sample question from the first to the last, sub-answer content respectively corresponding to the sample question under multiple thinking instructions, and the first probability of each sub-answer content to train the base model can enable the base model to learn the ability to split the problem step by step and answer it step by step. In this way, the ability of the base model to handle complex problems can be effectively improved. After training, even when faced with complex problems, the base model can analyze and answer them step by step, improving the accuracy of the output reply content.
[0120] In some embodiments, please refer to Figure 6 , Figure 6 FIG. further shows a schematic flowchart of the model training method provided by the embodiment of the present application. In step S110, the model training method may further include steps S510-S550:
[0121] S510. Obtain an initial data set, where the initial data set includes multiple reference questions and the reference reply content corresponding to each reference question.
[0122] Among them, the reference reply content corresponding to the reference question may be manually labeled or generated by the reference base model according to the reference question.
[0123] S520. The thinking chain generation model generates multiple thinking instructions for the reference question. The multiple thinking instructions for the reference question are used to represent multiple analysis steps for answering the reference question.
[0124] Among them, the chain-of-thought generation model is a pre-trained neural network model, whose input is a reference question and whose output is a series of step-by-step thinking instructions for the reference question.
[0125] In some embodiments, the chain-of-thought generation model can generate a linear chain of thought for the reference question, where the linear chain of thought contains a series of step-by-step thinking instructions for the reference question; it can also generate a tree-like chain of thought for the reference question, where the tree-like chain of thought includes multiple branches of thinking instruction steps, and each branch of thinking instruction steps contains a series of step-by-step thinking instructions. Subsequently, a search algorithm (such as beam search or Monte Carlo tree search) can be used to determine the optimal branch of thinking instruction steps from the tree-like chain of thought, and the series of step-by-step thinking instructions contained in the optimal branch of thinking instruction steps are used as the series of step-by-step thinking instructions for the reference question.
[0126] S530: The reference base model generates a response based on the reference question and the series of step-by-step thinking instructions for the reference question, obtaining candidate sub-response contents for the reference question under the series of step-by-step thinking instructions.
[0127] Among them, the reference base model can be the base model that needs to be trained mentioned in the foregoing embodiments, or it can be other base models.
[0128] S540: If the semantic similarity between the candidate sub-response content of the reference question under the corresponding last thinking instruction step and the reference response content corresponding to the reference question exceeds the similarity threshold, the reference question, the series of step-by-step thinking instructions for the reference question, and the candidate sub-response contents corresponding to the sample under the series of step-by-step thinking instructions are used as a training sample and added to the training sample set.
[0129] Among them, there are various calculation methods for the semantic similarity between the candidate sub-response content and the reference response content. For example, the edit distance between the response contents (including the candidate sub-response content and the reference response content) can be used as the semantic similarity between the response contents; or the response contents can be mapped to a vector space to obtain response vectors for the response contents, and the similarity between the response vectors can be used as the semantic similarity between the response contents. The specific method is not limited here.
[0130] Refer to the candidate sub-response content under the last thinking step instruction corresponding to the reference question, which is also the response content finally output by the reference base model for the reference question; it can be understood that if the semantic similarity between the candidate sub-response content under the last thinking step instruction corresponding to the reference question and the reference response content corresponding to the reference question exceeds the similarity threshold, it means that the reference base model has output the correct result for the reference question, and the reasoning process for the reference question (i.e., the candidate sub-response content under other thinking step instructions) is also probably correct. Therefore, the reference question, the multiple thinking step instructions of the reference question, and the candidate sub-response content corresponding to the sample under multiple thinking step instructions can be used as a training sample and added to the training sample set.
[0131] S550. If the semantic similarity between the candidate sub-response content under the last thinking step instruction corresponding to the reference question and the reference response content corresponding to the reference question does not exceed the similarity threshold, guide the reference base model to regenerate the response according to the reference question and the multiple thinking step instructions of the reference question.
[0132] Similarly, if the semantic similarity between the candidate sub-response content under the last thinking step instruction corresponding to the reference question and the reference response content corresponding to the reference question does not exceed the similarity threshold, it means that the reference base model has output an incorrect result for the reference question, and the reasoning process for the reference question (i.e., the candidate sub-response content under other thinking step instructions) is also probably incorrect. Therefore, the reference base model can be guided to regenerate the response according to the reference question and the multiple thinking step instructions of the reference question. Among them, guiding the reference base model to regenerate the response according to the reference question and the multiple thinking step instructions of the reference question means that, based on the prior knowledge that "the current candidate sub-response content is incorrect", the reference base model regenerates the response according to the reference question and the multiple thinking step instructions of the reference question, so as to avoid the reference base model from regenerating the same candidate sub-response content.
[0133] In some embodiments, the number of times each reference question is generated can also be set. After the number of times the reference question is generated reaches the threshold, the generation of a response for this reference question is abandoned, so as to avoid the reference base model falling into an infinite loop by repeatedly generating responses for the same reference question. Exemplarily, the initial value of the number of times each reference question is generated is set to 0, and the threshold is 2. When the semantic similarity between the candidate sub-response content of the reference question under the corresponding last thinking step instruction and the reference response content corresponding to the reference question does not exceed the similarity threshold, the number of times is incremented by 1. After that, if the number of times is less than the threshold of 2, the reference base model is guided to regenerate a response based on the reference question and multiple thinking step instructions of the reference question. If the number of times is not less than the threshold of 2, the generation of a response for this reference question is abandoned.
[0134] Of course, when the semantic similarity between the candidate sub-response content of the reference question under the corresponding last thinking step instruction and the reference response content corresponding to the reference question does not exceed the similarity threshold, the generation of a response for this reference question can also be directly abandoned.
[0135] Please refer to Figure 7 , after step S150, the method may further include steps S610 - S640:
[0136] S610. Obtain the target question.
[0137] Among them, the target question can be pure text content, or multi-modal content, such as the combination of image content and text content, the combination of audio content and text content, the combination of audio-visual content, etc.
[0138] It should be noted that the type of the target question needs to be consistent with the type of the sample questions in the training samples of the base model. For example, if the target question is pure text content, the sample questions used to train the base model also need to be pure text content; if the target question is the combination of image content and text content, the sample questions used to train the base model also need to be the combination of image content and text content.
[0139] S620. The thinking chain generation model generates multiple thinking step instructions for the target question based on the target question.
[0140] Similarly, multiple thinking step instructions for the target problem are used to represent multiple analysis steps for answering the target problem. Specifically, the target problem is input into the thinking chain generation model, and the thinking chain generation model outputs multiple thinking step instructions for the target problem; by automatically generating multiple thinking step instructions for the target problem from first to last by the thinking chain generation model, compared with manually writing thinking step instructions, the generation efficiency of thinking step instructions is greatly improved, and the quality of the generated thinking step instructions for different target problems is ensured, thereby improving the quality of the subsequent responses of the base model.
[0141] It can be understood that the number of thinking step instructions generated by the thinking chain generation model for different target problems, as well as the content of each thinking step instruction, are different and are related to the specific content of the target problem.
[0142] In some embodiments, multiple thinking step instructions for the target problem generated from first to last can be displayed in the final output target response content, so as to facilitate the user to understand the reasoning steps of the model; of course, they can also not be displayed in the final output target response content, allowing the user to only focus on the final target response content.
[0143] S630. The base model generates a response according to the target problem and multiple thinking step instructions for the target problem, and obtains the target sub-response content of the target problem under multiple thinking step instructions.
[0144] Among them, the base model is trained according to the method provided in the foregoing embodiments.
[0145] S640. Use the target sub-response content of the target problem under the corresponding last thinking step instruction as the target response content of the target problem.
[0146] In the prior art, the base model is usually guided to generate responses through a small number of examples with detailed reasoning steps, such that the quality of the responses generated by the base model largely depends on the quality and relevance of the selected examples; for questions that are quite different from the examples or relatively complex questions, it is difficult to guarantee the quality of the responses generated by the base model; alternatively, the base model is guided to generate responses by manually debugging the prompting words, but manually debugging the prompting words requires a lot of manpower and material resources to optimize the prompting words, and for questions in different fields, the optimization directions and optimization logics of the prompting words may not be the same, and everyone's criticism criteria are also affected by subjective consciousness, making the effect of guiding the base model to generate responses based on the prompting words unstable. In the above-described embodiments, for each target question, a plurality of thinking step instructions from the first to the last of the target question are generated by the thinking chain generation model, so that the base model generates a response according to the target question and the plurality of thinking step instructions of the target question to obtain the target response content; compared with the response generation process in the prior art, by generating a plurality of thinking step instructions from the first to the last of the target question, the base model can learn the thinking reasoning ability according to the plurality of thinking step instructions from the first to the last of each target question, thereby improving the question-and-answer effect of the base model. At the same time, by automatically generating a plurality of thinking step instructions from the first to the last of the target question by the thinking chain generation model, the generation efficiency of the thinking step instructions is improved, the time-consuming of manual annotation is reduced, and for target questions in different fields, the thinking chain generation model can ensure the consistency of the quality of the generated thinking step instructions, so that the base model can be applied to target questions in different fields.
[0147] In some embodiments, referring to Figure 8 , Figure 8 shows the flowchart of step S630 provided by the embodiment of the present application, and step S630 may include steps S710-S740: Figure 7
[0148] S710. The base model outputs the target sub-response content of the target question in the first thinking step instruction according to the target question and the first thinking step instruction for the target question.
[0149] S720. The base model predicts the target sub-response content of the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i-1)-th instruction response pair, and the i-th thinking step instruction of the target question; wherein, the (i-1)-th instruction response pair includes the (i-1)-th thinking step instruction for the target question and the target sub-response content of the target question under the (i-1)-th thinking step instruction; wherein, i is a positive integer greater than 1 and less than M, and M is the total number of thinking step instructions corresponding to the target question.
[0150] S730. Determine the accuracy score of the target sub-response content of the target problem in the i-th thinking step instruction.
[0151] In some embodiments, the accuracy score of the target sub-response content can be directly determined according to the probabilities of each character in the target sub-response content; for example, perform a weighted sum of the probabilities of each character in the target sub-response content, and use the sum result as the accuracy score of the target sub-response content.
[0152] In some other embodiments, an external discriminant model can also be called, and the target problem, multiple thinking step instructions of the target problem from the first to the last, and the target sub-response content are input into the discriminant model, and the discriminant model outputs the accuracy of the target sub-response content.
[0153] Among them, the discriminant model is pre-trained. When training the discriminant model, the input can be a discriminant problem, multiple thinking step instructions of the discriminant problem from the first to the last, the sub-response content of the discriminant problem under any thinking step instruction, and an accuracy label, and the output is a predicted accuracy score. According to the predicted accuracy score and the accuracy label, calculate the model loss of the discriminant model. Then, adjust the parameters of the discriminant model according to the model loss until the training end condition of the discriminant model is reached.
[0154] S740. If the accuracy score of the target sub-response content of the target problem in the i-th thinking step instruction is greater than the score threshold, increment i by 1, and return to execute the target sub-response content predicted by the base model for the target problem under the i-th thinking step instruction according to the target problem, the first i - 1 instruction response pairs of the target problem, and the i-th thinking step instruction of the target problem.
[0155] The accuracy score of the target sub-response content of the i-th thinking step instruction being greater than the score threshold indicates that the target sub-response content of the i-th thinking step instruction is accurate. At this time, the corresponding sub-response content for the next thinking step instruction can be continued to be generated.
[0156] In the above embodiments, after generating the target sub-response content corresponding to each thinking step instruction, calculate the accuracy score of the target sub-response content of this thinking step instruction, and only continue with the subsequent response generation on the basis that the target sub-response content is accurate enough, realizing real-time detection of the response generation process of the base model, ensuring the accuracy of the target sub-response content corresponding to each thinking step instruction, avoiding the spread of incorrect responses, and thus ensuring the reliability of the target response content finally output by the base model.
[0157] In some embodiments, please refer to Figure 9 , Figure 9Another flowchart of the reply generation method provided by the embodiment of the present application is shown. After step S730, the reply generation method may further include steps S810 - S830:
[0158] S810. If the determination score of the target sub - reply content of the target problem in the i - th thinking step instruction is not greater than the score threshold, increment the reply count corresponding to the i - th thinking step instruction by 1.
[0159] The accuracy score of the target sub - reply content of the i - th thinking step instruction not being greater than the score threshold indicates that the target sub - reply content of the i - th thinking step instruction is not accurate enough. Since the sub - reply content corresponding to the next thinking step instruction needs to be generated based on the sub - reply content corresponding to the current thinking step instruction, if the reply generation continues for the next thinking step instruction, the obtained sub - reply content is also likely to be inaccurate, resulting in an inaccurate final target reply content, that is, an invalid final target reply content is output. Therefore, when the determination score of the target sub - reply content of the i - th thinking step instruction is not greater than the score threshold, the reply count corresponding to the i - th thinking step instruction can be incremented by 1, and when the reply count corresponding to this thinking step instruction is less than the reply count threshold, the base model is guided to regenerate the target sub - reply content of this thinking step instruction.
[0160] S820. If the reply count corresponding to the i - th thinking step instruction is less than the reply count threshold, guide the base model to regenerate the target sub - reply content for the target problem under the i - th thinking step instruction according to the target problem, the target problem from the 1st instruction - reply pair to the (i - 1) - th instruction - reply pair, and the i - th thinking step instruction of the target problem.
[0161] It should be noted that for the 1st thinking step instruction, if the reply count corresponding to the 1st thinking step instruction is less than the reply count threshold, guide the base model to re - output the target sub - reply content of the 1st thinking step instruction according to the target problem and the 1st thinking step instruction of the target problem.
[0162] S830. If the reply count corresponding to the i - th thinking step instruction is not less than the reply count threshold, increment i by 1, and return to execute the prediction by the base model of the target sub - reply content for the target problem under the i - th thinking step instruction according to the target problem, the target problem from the 1st instruction - reply pair to the (i - 1) - th instruction - reply pair, and the i - th thinking step instruction of the target problem.
[0163] In the above embodiments, when the determination score of the target sub-response content of the i-th thinking step instruction is not greater than the score threshold, the base model is guided to regenerate the target sub-response content of the thinking step instruction, so as to ensure that the target sub-response content of each thinking step instruction is accurate enough. At the same time, by introducing the number of responses corresponding to each thinking step instruction, the situation that the base model falls into an infinite loop due to its inability to generate target sub-response content with an accuracy score greater than the score threshold is avoided, ensuring the normal operation of the response generation process of the base model.
[0164] In some embodiments, please refer to Figure 10 , Figure 10 which shows Figure 7 the flowchart of step S620 provided by the embodiment of the present application. Step S620 may include steps S910 - S930:
[0165] S910. The thinking chain model generates thinking step instructions based on the target problem to obtain a tree-shaped thinking chain of the target problem. A tree-shaped thinking chain of a target problem includes multiple thinking step instruction branches, and each thinking step instruction branch contains multiple thinking step instructions of the target problem.
[0166] S920. Calculate the path score of each thinking step instruction branch, and based on the path scores of the thinking step instruction branches of each target problem, determine the target thinking step instruction branch from the multiple thinking step instruction branches included in the tree-shaped thinking chain of the target problem.
[0167] Among them, the path score of each thinking step instruction branch can be calculated by a path search algorithm; the path search algorithm is, for example, beam search, Monte Carlo tree search, etc.
[0168] S930. Use the multiple thinking step instructions included in the target thinking step instruction branch as the multiple thinking step instructions of the target problem.
[0169] In the above embodiments, since the tree-shaped thinking chain includes multiple possible thinking step instruction branches, by selecting the target thinking step instruction branch with the highest path score from the multiple thinking step instruction branches, the multiple thinking step instructions of the target problem input to the base model from the first to the last are the optimal thinking step instructions, thereby improving the ability of the base model to handle complex and uncertain problems.
[0170] In some embodiments, please refer to Figure 11 , Figure 11 which shows the schematic diagram of the model training device provided by the embodiment of the present application. The model training device 1000 includes:
[0171] The first acquisition module 1010 is configured to acquire a training sample set. Each training sample in the training sample set includes a sample question, a plurality of thinking step instructions, a plurality of sub-response contents, and a first probability of each of the sub-response contents. The plurality of thinking step instructions are used to represent a plurality of analysis steps for answering the sample question. The plurality of thinking step instructions correspond to the plurality of sub-response contents one by one. One sub-response content is a response content obtained by analyzing the sample question in the analysis step represented by one thinking step instruction.
[0172] The prediction module 1020 is configured to, according to the sample question and the plurality of thinking step instructions of the sample question, sequentially predict, by the base model, the prediction probability of outputting the corresponding sub-response content under each thinking step instruction of the sample question.
[0173] The first loss calculation module 1030 is configured to determine a response generation loss based on the prediction probabilities of the sub-response contents respectively corresponding to the sample question under the plurality of thinking step instructions.
[0174] The second loss calculation module 1040 is configured to determine a parameter update loss according to the first probability and the prediction probability of the sub-response contents respectively corresponding to the sample question under the plurality of thinking step instructions.
[0175] The parameter adjustment module 1050 is configured to update the parameters of the base model according to the response generation loss and the parameter update loss until the training end condition is reached.
[0176] In some embodiments, the second loss calculation module 1040 includes a reward module configured to predict, by the value network, the expected cumulative reward corresponding to the sub-response content of the sample question under the t-th thinking step instruction according to the sub-response content of the sample question under the t-th thinking step instruction, where t is a positive integer starting from 1, t ≤ N, and N is the total number of thinking step instructions corresponding to the sample question; an advantage calculation module configured to estimate the advantage function of the sub-response content of the sample question under the t-th thinking step instruction according to the expected cumulative reward corresponding to the sub-response content of the sample question under the t-th thinking step instruction and the prediction probability of the sub-response content of the sample question under the t-th thinking step instruction; a probability ratio calculation module configured to determine the probability ratio of the sub-response content of the sample question under the t-th thinking step instruction according to the first probability and the prediction probability of the sub-response content of the sample question under the t-th thinking step instruction; and an expectation calculation module configured to calculate the expectation of the base model sequentially outputting the sub-response contents of the sample question from the 1st thinking step instruction to the Nth thinking step instruction according to the probability ratios and the advantage functions of the sub-response contents of the sample question from the 1st thinking step instruction to the Nth thinking step instruction, as the parameter update loss.
[0177] In some embodiments, the desired computing module includes a first advantage return computing module for calculating the product of the probability ratio of the sub-response content of the sample problem at the t-th thinking step instruction and the corresponding advantage function to obtain the first advantage return of the sub-response content of the sample problem at the t-th thinking step instruction; a clipping module for clipping the probability ratio of the sub-response content of the sample problem at the t-th thinking step instruction according to a preset update amplitude interval to obtain the target probability ratio of the sub-response content of the sample problem at the t-th thinking step instruction; the target probability ratio is within the update amplitude interval; a second advantage return computing module for calculating the product of the advantage function of the sub-response content of the sample problem at the t-th thinking step instruction and the corresponding target probability ratio to obtain the second advantage return of the sub-response content of the sample problem at the t-th thinking step instruction; a selection module for taking the minimum value of the first advantage return and the second advantage return of the sub-response content of the sample problem at the t-th thinking step instruction as the target advantage return of the sub-response content of the sample problem at the t-th thinking step instruction; an accumulation module for calculating the expectation of the target advantage return of the sub-response content of the sample problem from the first thinking step instruction to the N-th thinking step instruction to obtain the parameter update loss.
[0178] In some embodiments, when t is greater than 1, the base model predicts the prediction probability of the corresponding sub-response content of the sample problem at the t-th thinking step instruction according to the following process: the base model sequentially predicts the prediction probabilities of each character in the corresponding sub-response content of the sample problem at the t-th thinking step instruction according to the sample problem, the sample problem from the first instruction response pair to the (t - 1)-th instruction response pair, and the t-th thinking step instruction of the sample problem; the (t - 1)-th instruction response pair includes the (t - 1)-th thinking step instruction of the sample problem and the sub-response content of the sample problem at the (t - 1)-th thinking step instruction; according to the prediction probabilities of each character in the corresponding sub-response content of the sample problem at the t-th thinking step instruction, the prediction probability of the corresponding sub-response content of the sample problem at the t-th thinking step instruction is determined; wherein, the base model sequentially predicts the prediction probabilities of each character in the corresponding sub-response content of the sample problem at the first thinking step instruction according to the sample problem and the first thinking step instruction of the sample problem, and determines the prediction probability of the corresponding sub-response content of the sample problem at the first thinking step instruction.
[0179] In some embodiments, the model training device 1000 further includes a data preparation module for obtaining an initial data set, where the initial data set includes a plurality of reference questions and corresponding reference response contents for each reference question; an instruction generation module for generating a plurality of thinking step instructions for the reference questions by a thinking chain generation model; the plurality of thinking step instructions for the reference questions are used to represent a plurality of analysis steps for answering the reference questions; a reference response module for generating a response by a reference base model according to the reference question and the plurality of thinking step instructions for the reference question, and obtaining candidate sub-response contents for the reference question under the plurality of thinking step instructions; an adding module for adding, as a training sample, the reference question, the plurality of thinking step instructions for the reference question, and the candidate sub-response contents respectively corresponding to the sample under the plurality of thinking step instructions to the training sample set if the semantic similarity between the candidate sub-response content of the reference question under the corresponding last thinking step instruction and the reference response content corresponding to the reference question exceeds a similarity threshold.
[0180] In some embodiments, the model training device 1000 further includes an execution module for guiding the reference base model to regenerate a response according to the reference question and the plurality of thinking step instructions for the reference question if the semantic similarity between the candidate sub-response content of the reference question under the corresponding last thinking step instruction and the reference response content corresponding to the reference question does not exceed the similarity threshold.
[0181] In some embodiments, the model training device further includes:
[0182] A second acquisition module for acquiring a target question.
[0183] A thinking chain generation module for generating a plurality of thinking step instructions for the target question by a thinking chain generation model based on the target question.
[0184] A reply module for generating a reply by a base model according to the target question and the plurality of thinking step instructions for the target question, and obtaining target sub-reply contents for the target question under the plurality of thinking step instructions.
[0185] An output module for using the target sub-reply content of the target question under the corresponding last thinking step instruction as the target reply content of the target question.
[0186] In some embodiments, the response module is specifically configured to: have the base model output the target sub-response content of the target question under the first thinking step instruction according to the target question and the first thinking step instruction for the target question; have the base model sequentially predict the target sub-response content output for the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction for the target question; wherein, the (i - 1)-th instruction response pair includes the (i - 1)-th thinking step instruction for the target question and the target sub-response content of the target question under the (i - 1)-th thinking step instruction; wherein, i is a positive integer greater than 1 and less than M, and M is the total number of thinking step instructions corresponding to the target question; determine the accuracy score of the target sub-response content of the target question under the i-th thinking step instruction; if the accuracy score of the target sub-response content of the target question under the i-th thinking step instruction is greater than the score threshold, increment i by 1, and return to execute having the base model predict the target sub-response content output for the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction for the target question.
[0187] In some embodiments, the response module is further configured to: if the determination score of the target sub-response content of the target question under the i-th thinking step instruction is not greater than the score threshold, increment the number of response times corresponding to the i-th thinking step instruction by 1; if the number of response times corresponding to the i-th thinking step instruction is less than the response times threshold, guide the base model to regenerate the target sub-response content output for the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction for the target question; if the number of response times corresponding to the i-th thinking step instruction is not less than the response times threshold, increment i by 1, and return to execute having the base model predict the target sub-response content output for the target question under the i-th thinking step instruction according to the target question, the target question from the first instruction response pair to the (i - 1)-th instruction response pair, and the i-th thinking step instruction for the target question.
[0188] In some embodiments, the thought chain generation module includes an expansion module for generating thought step instructions according to a target question by a thought chain model to obtain a tree-shaped thought chain of the target question. A tree-shaped thought chain of a target question includes multiple thought step instruction branches, and each thought step instruction branch contains multiple thought step instructions of the target question from the first to the last; a screening module for calculating the path scores of each thought step instruction branch and determining a target thought step instruction branch from the multiple thought step instruction branches included in the tree-shaped thought chain of the target question based on the path scores of the thought step instruction branches of each target question; and a determination module for using the multiple thought step instructions included in the target thought step instruction branch as the multiple thought step instructions of the target question.
[0189] Figure 12 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. The electronic device may be the above terminal for implementing the model training method and the response generation method provided by the present application. It should be noted that Figure 12 The computer system 1300 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0190] As Figure 12 shown, the computer system 1300 includes a central processing unit (CPU)
[0191] 1301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1302 or the program loaded from the storage section 1308 into the random access memory (RAM) 1303, such as executing the methods in the above embodiments. In the RAM 1303, various programs and data required for system operations are also stored. The CPU 1301, the ROM 1302, and the RAM 1303 are connected to each other through a bus 1304. The input / output (I / O) interface 1305 is also connected to the bus 1304.
[0192] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, a microphone, etc.; an output section 1307 including such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. The drive 1310 is also connected to the I / O interface 1305 as required. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as required so that computer instructions read from it are installed into the storage section 1308 as required.
[0193] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes computer instructions. When the computer instructions are executed by a central processing unit (CPU) 1301, various functions defined in the system of the present application are executed.
[0194] The present application also provides a computer-readable storage medium storing computer instructions, which when executed by a processor, implement the method in any of the above method embodiments.
[0195] It should be noted that the computer-readable storage medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0196] In the embodiments of the present application, the term "module" or "unit" refers to a computer instruction with a predetermined function or a part of a computer instruction, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be an overall module or a part of the unit of the function of the module or unit.
[0197] The above are only the preferred embodiments of the present application, and do not impose any formal restrictions on the present application. Although the present application has been disclosed above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to equivalent embodiments of equivalent changes by using the above-disclosed technical content within the scope of the technical solution of the present application. However, as long as it does not depart from the content of the technical solution of the present application, any brief modifications, equivalent changes and modifications made to the above embodiments according to the technical essence of the present application still fall within the scope of the technical solution of the present application.
Claims
1. A model training method, characterized in that: include: Acquire a training sample set, wherein each training sample in the training sample set includes a sample question, a plurality of thinking step instructions, a plurality of sub-reply contents, and a first probability of each of the sub-reply contents, wherein the plurality of thinking step instructions are used to represent a plurality of analysis steps for answering the sample question, the plurality of thinking step instructions correspond one-to-one to the plurality of sub-reply contents, and one sub-reply content is a reply content obtained by analyzing the sample question at an analysis step represented by a thinking step instruction; The base model sequentially predicts the predicted probability of outputting the corresponding sub-response content under each thinking step instruction of the sample question according to the sample question and the multiple thinking step instructions of the sample question; Determining a response generation loss based on predicted probabilities of sub-response contents corresponding to the sample question under the plurality of thought step instructions; Determine the parameter update loss according to the first probability and the predicted probability of the sub-response content corresponding to the sample question under the plurality of thinking step instructions; The parameters of the base model are updated according to the response generation loss and the parameter update loss until the training end condition is reached.
2. The method according to claim 1, characterized in that The determining of the parameter update loss according to the first probability and the predicted probability of the sub-reply contents respectively corresponding to the sample question under the plurality of thinking step instructions comprises: The value network predicts the expected cumulative reward corresponding to the sub-reply content under the t-th thinking step instruction of the sample question based on the sub-reply content under the t-th thinking step instruction of the sample question; t is a positive integer starting from 1, t≤N, and N is the total number of thinking step instructions corresponding to the sample question; Estimate the advantage function of the sub-reply content under the t-th thinking step instruction for the sample question according to the expected cumulative reward corresponding to the sub-reply content under the t-th thinking step instruction for the sample question and the predicted probability of the sub-reply content under the t-th thinking step instruction for the sample question; Determine a probability ratio of the sub-reply content under the t-th thinking step instruction for the sample question according to the first probability and the predicted probability of the sub-reply content under the t-th thinking step instruction for the sample question; Based on the probability ratio and advantage function of the sub-reply contents under the 1st thinking step instruction to the Nth thinking step instruction for the sample problem, the base model is calculated to sequentially output the expectation of the sub-reply contents under the 1st thinking step instruction to the Nth thinking step instruction for the sample problem as the parameter update loss.
3. The method according to claim 2, characterized in that The method of calculating the expectation of the sub-reply contents under the first thinking step instruction to the Nth thinking step instruction sequentially outputted by the base model for the sample problem as the parameter update loss according to the probability ratio and advantage function of the sub-reply contents under the first thinking step instruction to the Nth thinking step instruction for the sample problem comprises: Calculate the product of the probability ratio of the sub-reply content of the t-th thinking step instruction for the sample question and the corresponding advantage function to obtain the first advantage return of the sub-reply content of the t-th thinking step instruction for the sample question; According to a preset update amplitude interval, a probability ratio of the sub-reply content of the sample question at the t-th thinking step instruction is clipped to obtain a target probability ratio of the sub-reply content of the sample question at the t-th thinking step instruction; the target probability ratio is within the update amplitude interval; Calculate the product of the advantage function of the sub-reply content of the t-th thinking step instruction for the sample question and the corresponding target probability ratio to obtain a second advantage return of the sub-reply content of the t-th thinking step instruction for the sample question; The minimum value between the first advantage return and the second advantage return of the sub-response content of the t-th thinking step instruction for the sample question is used as the target advantage return of the sub-response content of the t-th thinking step instruction for the sample question; Calculate the expectation of the target advantage return of the sub-response content from the 1st thinking step instruction to the target advantage return of the sub-response content of the Nth thinking step instruction for the sample problem, and obtain the parameter update loss.
4. The method according to claim 1, characterized in that The base model sequentially predicts the predicted probability of outputting the corresponding sub-reply content under each thinking step instruction of the sample question according to the sample question and the multiple thinking step instructions of the sample question, including: When t is greater than 1, the base model predicts the predicted probability of outputting the corresponding sub-response content under the t-th thinking step instruction of the sample question according to the following process: The base model sequentially predicts the predicted probability of each character in the sub-reply content corresponding to the sample question under the t-th thinking step instruction according to the sample question, the sample question from the 1st instruction-reply pair to the t-1th instruction-reply pair and the t-th thinking step instruction of the sample question; the t-1th instruction-reply pair includes the t-1th thinking step instruction for the sample question and the sub-reply content for the sample question under the t-1th thinking step instruction; Determine the predicted probability of the sub-reply content corresponding to the sample question output under the t-th thinking step instruction according to the predicted probability of each character in the sub-reply content corresponding to the sample question output under the t-th thinking step instruction; Among them, the base model sequentially predicts the predicted probability of each character in the sub-reply content corresponding to the sample question under the first thinking step instruction based on the sample question and the first thinking step instruction for the sample question, and determines the predicted probability of outputting the corresponding sub-reply content for the sample question under the first thinking step instruction.
5. The method according to claim 1, characterized in that: Before obtaining the training sample set, the method further includes: Acquire an initial data set, the initial data set including a plurality of reference questions and reference answer content corresponding to each of the reference questions; Generate a plurality of thinking step instructions for the reference problem from a thinking chain generation model; the plurality of thinking step instructions for the reference problem are used to represent a plurality of analysis steps for solving the reference problem; The reference base model generates a response according to the reference question and multiple thinking step instructions of the reference question, and obtains candidate sub-response contents of the reference question under the multiple thinking step instructions; If the semantic similarity between the candidate sub-reply content under the corresponding last thinking step instruction of the reference question and the reference reply content corresponding to the reference question exceeds a similarity threshold, the reference question, multiple thinking step instructions of the reference question and the candidate sub-reply content corresponding to the sample under the multiple thinking step instructions are taken as a training sample and added to the training sample set.
6. The method according to claim 5, characterized in that The method further comprises: If the semantic similarity between the candidate sub-reply content under the corresponding last thinking step instruction of the reference question and the reference reply content corresponding to the reference question does not exceed the similarity threshold, guide the reference base model to re-generate a reply based on the reference question and multiple thinking step instructions of the reference question.
7. The method according to any one of claims 1 to 6, characterized in that After the parameters of the base model are updated according to the reply generation loss and the parameter update loss until the training end condition is reached, the method further comprises: Get the target question; The thinking chain generation model generates a plurality of thinking step instructions for the target problem based on the target problem; The base model generates a response according to the target question and multiple thinking step instructions of the target question, and obtains target sub-response content of the target question under the multiple thinking step instructions; The target sub-reply content of the target question under the corresponding last thinking step instruction is used as the target answer content of the target question.
8. The method according to claim 7, characterized in that The base model generates a response according to the target question and the multiple thinking step instructions of the target question to obtain the target sub-response content of the target question under the multiple thinking step instructions, including: The base model outputs the target sub-response content of the target problem in the first thinking step instruction according to the target problem and the first thinking step instruction for the target problem; The base model predicts the target sub-reply content of the target problem under the i-th thinking step instruction according to the target problem, the target problem from the 1st instruction-reply pair to the i-1th instruction-reply pair and the i-th thinking step instruction of the target problem; wherein the i-1th instruction-reply pair includes the i-1th thinking step instruction for the target problem and the target sub-reply content for the target problem under the i-1th thinking step instruction; wherein i is a positive integer greater than 1 and less than M, and M is the total number of thinking step instructions corresponding to the target problem; Determining the accuracy score of the target sub-response content of the target question in the i-th thinking step instruction; If the accuracy score of the target sub-reply content of the target question in the i-th thinking step instruction is greater than the score threshold, i is increased by 1, and the base model is returned to execute, based on the target question, the target question from the 1st instruction-reply pair to the i-1th instruction-reply pair and the i-th thinking step instruction of the target question, to predict the target sub-reply content for the target question under the i-th thinking step instruction.
9. The method according to claim 8, characterized in that After determining the accuracy score of the target sub-response content of the i-th thinking step instruction of the target question, the method further includes: If the accuracy score of the target sub-response content of the target question in the i-th thinking step instruction is not greater than the score threshold, the number of replies corresponding to the i-th thinking step instruction is increased by 1; If the number of replies corresponding to the i-th thinking step instruction is less than the reply number threshold, guide the base model to regenerate the target sub-reply content output under the i-th thinking step instruction for the target problem according to the target problem, the target problem from the 1st instruction-reply pair to the i-1th instruction-reply pair and the i-th thinking step instruction of the target problem; If the number of replies corresponding to the i-th thinking step instruction is not less than the reply number threshold, i is increased by 1, and the base model is returned to execute, based on the target problem, the target problem from the 1st instruction-reply pair to the i-1th instruction-reply pair and the i-th thinking step instruction of the target problem, to predict the target sub-reply content for the target problem under the i-th thinking step instruction.
10. The method according to claim 9, characterized in that The thinking chain model generates thinking step instructions according to the target problem to obtain multiple thinking step instructions for the target problem, including: The thinking chain model generates thinking step instructions according to the target problem to obtain a tree-shaped thinking chain of the target problem, wherein a tree-shaped thinking chain of the target problem includes multiple thinking step instruction branches, and each of the thinking step instruction branches includes multiple thinking step instructions of the target problem from the earliest to the latest. Calculating the path score of each of the thought step instruction branches, and based on the path score of each of the thought step instruction branches of the target problem, determining a target thought step instruction branch from a plurality of thought step instruction branches included in the tree-like thought chain of the target problem; The multiple thinking step instructions included in the target thinking step instruction section are used as the multiple thinking step instructions for the target problem.
11. A model training device, characterized in that: include: A first acquisition module is used to acquire a training sample set, each training sample in the training sample set includes a sample question, a plurality of thinking step instructions, a plurality of sub-reply contents, and a first probability of each of the sub-reply contents, the plurality of thinking step instructions are used to represent a plurality of analysis steps for answering the sample question, the plurality of thinking step instructions correspond to the plurality of sub-reply contents one by one, and one sub-reply content is a reply content obtained by analyzing the sample question at an analysis step represented by a thinking step instruction; A prediction module, configured to use the base model to sequentially predict the predicted probability of outputting the corresponding sub-reply content under each thinking step instruction of the sample question according to the sample question and the multiple thinking step instructions of the sample question; A first loss calculation module, configured to determine a reply generation loss based on the predicted probabilities of the sub-reply contents corresponding to the sample question under the plurality of thought step instructions; A second loss calculation module is used to determine the parameter update loss according to the first probability and the predicted probability of the sub-reply content corresponding to the sample question under the plurality of thinking step instructions; A parameter adjustment module is used to update the parameters of the base model according to the response generation loss and the parameter update loss until the training end condition is reached.
12. An electronic device, characterized in that: include: processor; A memory, wherein computer instructions are stored in the memory, and when the computer instructions are executed by the processor, the method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method described in any one of claims 1 to 10 is implemented.