Large model training method, question and answer method, related equipment and program product
By introducing exploration rewards and phased training in large model training, the problem of models tending to repeat known paths in traditional methods is solved, and the accuracy and efficiency of large models in complex reasoning tasks are improved.
Patent Information
- Application Number
- CN202510973166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional reinforcement learning methods lack an effective exploration incentive mechanism in large model training, which causes the model to tend to reuse known reasoning paths and ignore new paths, making it difficult to find better solutions in complex reasoning tasks. Insufficient feedback also leads to a lack of effective guidance in the intermediate steps of the model.
Exploration rewards are generated by calculating the hidden state features based on the large model inference process. Combined with the result rewards, reinforcement learning is used to update model parameters. Training is carried out in stages to encourage the model to explore unknown paths and provide immediate feedback.
It improves the performance of large models on complex reasoning tasks, improves the accuracy of answers to complex questions and the efficiency of reasoning, avoids local optimality, and increases the probability of the model finding a better reasoning path when facing complex problems.
Smart Images

Figure CN120806159A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large model training, and more particularly, to a large model training method, a question and answer method, related equipment and a computer program product. BACKGROUND
[0002] Reinforcement learning is a method of learning optimal behavior policy by interacting with the environment. In complex reasoning tasks such as mathematical problem solving, logical reasoning, etc., reinforcement learning is used to train large language models (LLM) to generate high-quality reasoning paths.
[0003] Traditional reinforcement learning methods generally calculate reward signals based on the results output by large models. This reward mechanism gives the same reward to all reasoning paths that lead to the correct answer, without distinguishing the quality of the reasoning paths, and thus tends to cause large models to rely on known reasoning paths and ignore the exploration of new reasoning paths. For example, if a large model finds a correct reasoning path for a certain problem, it will tend to reuse this path rather than trying a new reasoning path. This makes it difficult for the model to discover better reasoning paths when facing complex problems, reducing the processing effect of the large model on complex reasoning tasks, and the answers given to complex reasoning problems are not accurate enough. SUMMARY
[0004] In view of the above problems, the present application is proposed to provide a large model training method, a question and answer method, related equipment and a computer program product to encourage the model to explore new reasoning paths during large model training, and improve the processing effect of the trained large model on complex reasoning tasks. The specific scheme is as follows:
[0005] In a first aspect, a large model training method is provided, comprising:
[0006] Obtaining question and answer training data, the question and answer training data comprising problem samples and corresponding answer labels;
[0007] Sending the problem samples into a large model to be trained for reasoning to obtain a prediction output of the large model;
[0008] Calculating an exploration reward based on the hidden layer state features generated during the reasoning process of the large model, the exploration reward being used to encourage the large model to use unknown reasoning paths to process the problem samples;
[0009] Calculating a result reward based on the prediction output and the answer label;
[0010] Updating the parameters of the large model in a reinforcement learning manner according to the exploration reward and the result reward.
[0011] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of calculating the exploration reward based on the hidden layer state features generated by the large model inference process includes:
[0012] The hidden state features generated by the large model inference process are fed into the configured prediction network to obtain the probability distribution of the prediction results;
[0013] calculating a first statistic based on the probability distribution, the first statistic representing uncertainty of the prediction network;
[0014] An exploration reward is determined based on the first statistic, and a size of the exploration reward is positively correlated with the first statistic.
[0015] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of calculating the exploration reward based on the hidden layer state features generated by the large model inference process further includes:
[0016] calculating a second statistic based on the probability distribution, wherein the second statistic represents a prediction result;
[0017] Calculating a prediction loss with the goal of making the second statistic approach the answer label, and calculating a total loss based on the prediction loss and the exploration reward, where the total loss is positively correlated with the prediction loss and negatively correlated with the exploration reward;
[0018] The parameters of the prediction network are updated according to the total loss.
[0019] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of calculating the total loss based on the predicted loss and the exploration reward includes:
[0020] Calculating a total loss based on the prediction loss, the exploration reward, and an adjustable weight factor, wherein the weight factor is used to balance the weights of the prediction loss and the exploration reward in the total loss;
[0021] The large model training process includes at least two stages. By adjusting the size of the weight factor, the weight of the exploration reward in the second stage is smaller than the weight of the exploration reward in the first stage. The second stage is later than the first stage in time sequence.
[0022] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the first statistic is the variance of the probability distribution.
[0023] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the second statistic is the mean of the probability distribution.
[0024] In a possible design, in another implementation manner of the first aspect of the embodiment of the present application, after the exploration reward is calculated, the method further includes:
[0025] The exploration reward is decayed according to the current training step, to obtain a decayed exploration reward, and the decay degree of the exploration reward becomes larger with the increase of the training step;
[0026] Then, the process of updating the parameters of the large model in the reinforcement learning manner according to the exploration reward and the result reward includes:
[0027] The parameters of the large model are updated in the reinforcement learning manner according to the decayed exploration reward and the result reward.
[0028] In a second aspect, a large model training method is provided, including a training process of three stages of fast thinking stage, verification stage and slow thinking stage, wherein
[0029] In the fast thinking stage, the problem sample is taken as the input of the large model, the large model is instructed to generate an initial answer quickly, and a first reward is calculated based on the initial answer and a target answer corresponding to the problem sample;
[0030] In the verification stage, the initial answer output by the large model in the fast thinking stage is taken as the input of the large model, the large model is instructed to verify the correctness of the initial answer, a second reward is calculated according to the verification result output by the large model, and in the case that the verification result indicates that the initial answer is incorrect, the slow thinking stage is entered;
[0031] In the slow thinking stage:
[0032] The problem sample and the initial answer output by the large model in the fast thinking stage are taken as an input sequence of the large model, and the large model is instructed to infer and generate a revised answer;
[0033] An exploration reward is calculated based on the hidden layer state feature generated in the inference process of the large model, and the exploration reward is used to encourage the large model to process the input sequence by using an unknown inference path;
[0034] A result reward is calculated based on the revised answer and the target answer;
[0035] The parameters of the large model are updated in the reinforcement learning manner according to the exploration reward and the result reward.
[0036] In a third aspect, a large model-based question answering method is provided, including:
[0037] A problem description is obtained;
[0038] The question description is input into a configured large model to obtain a large model output answer result, wherein the large model is trained by the large model training method described in the first aspect or the second aspect.
[0039] In a fourth aspect, an electronic device is provided, comprising a memory and a processor.
[0040] The memory is configured to store a program.
[0041] The processor is configured to execute the program to implement the steps of the large model training method described in the first aspect or the second aspect, or to implement the steps of the large model-based question and answer method described in the third aspect.
[0042] In a fifth aspect, a readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the large model training method described in the first aspect or the second aspect are implemented, or the steps of the large model-based question and answer method described in the third aspect are implemented.
[0043] In a sixth aspect, a computer program product is provided, comprising a computer program. When the computer program is executed by a processor, the steps of the large model training method described in the first aspect or the second aspect are implemented, or the steps of the large model-based question and answer method described in the third aspect are implemented.
[0044] By the above technical solution, the large model is trained by using question and answer training data. On the basis of the prediction output of the large model and the answer label calculation result reward, an exploration reward is further calculated based on the hidden layer state features generated in the inference process of the large model. The exploration reward is used to encourage the large model to process the input question sample by using unknown inference paths. By referring to the exploration reward and the result reward, the parameters of the large model are updated in a reinforcement learning manner. As can be seen, the exploration reward is additionally added in the reinforcement learning process, which can encourage the large model to explore unknown inference paths and avoid falling into local optimum, thereby improving the probability of the large model finding better inference paths when facing complex problems, and further improving the performance of the large model on complex inference tasks and the accuracy of the answer result to complex inference problems.
[0045] Further, since the exploration reward is calculated based on the hidden layer state features generated in the inference process of the large model, the exploration reward can be calculated based on the hidden layer state features generated at different time points in the inference process of the large model. The exploration rewards at different time points can guide the large model to better adjust the inference path in the middle step of inference, that is, the step-by-step path optimization can be realized, and the inference performance of the large model is further improved, especially on complex inference tasks. BRIEF DESCRIPTION OF DRAWINGS
[0046] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments, and are not intended to limit the scope of the present application. Furthermore, like reference numerals are intended to represent the same components throughout the various figures. In the drawings:
[0047] Figure 1 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided;
[0048] Figure 2 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided;
[0049] Figure 3 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided;
[0050] Figure 4 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided;
[0051] Figure 5 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided;
[0052] Figure 6 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided;
[0053] Figure 7 An embodiment system architecture schematic diagram of a large model training method provided for an embodiment of the present application is provided. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0055] The large model training process generally adopts a reinforcement learning training method, such as a proximal policy optimization (PPO) algorithm, a group regularization policy optimization (GRPO) algorithm, etc.
[0056] The traditional reinforcement learning algorithm simply calculates the reward signal based on the result, and is prone to the following limitations:
[0057] 1. Lack of an effective exploration incentive mechanism: Training strategies that calculate reward signals solely based on results can easily lead models to favor known reasoning paths while neglecting the exploration of new reasoning paths. This is primarily because the result-based reward signal gives the same reward to all paths leading to the correct answer, without distinguishing between the quality of the paths. For example, if the model finds a correct reasoning path for a problem, it will tend to reuse this path rather than try new ones. Traditional methods lack effective mechanisms to incentivize models to explore new reasoning paths. This makes it difficult for the model to discover a better solution path when faced with complex problems. For example, when solving complex logical reasoning problems, the model may reuse known reasoning paths and ignore potentially more effective paths.
[0058] 2. Insufficient feedback during multi-step reasoning: In complex reasoning tasks, reward signals are often sparse, meaning they are only given when the final answer is correct. This sparsity results in a lack of effective feedback during intermediate reasoning steps, making it difficult for the model to learn the correct reasoning path. For example, in mathematical problem solving, rewards are only given when the final answer is correct. During the intermediate steps, the model cannot obtain feedback on the correctness of its reasoning path, making it difficult for the model to gradually optimize its reasoning path.
[0059] In view of this, an embodiment of the present application provides a large model training solution that can solve at least some of the defects of the existing technology.
[0060] The present invention provides a large model training method, which can be applied to Figure 1 The system architecture shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 (This section includes a server as an example).
[0061] The terminal 100 or the server 200 can be used alone to execute the large model training method provided in the embodiment of the present application. In addition, the terminal 100 and the server 200 can also be used in conjunction to execute the large model training method provided in the embodiment of the present application.
[0062] Next describe Figure 1 The product form of the mid-terminal 100;
[0063] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a teaching screen, a wearable device, a vehicle-mounted device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0064] The embodiment of the present application provides a large model training method, which is illustrated by applying the method to a computer device. The computer device can be Figure 1 The terminal 100 or the system consisting of the terminal 100 and the server 200. Figure 2 , the large model training method specifically includes the following steps:
[0065] Step S100: Obtain question-answering training data, which includes question samples and corresponding answer labels.
[0066] The large model training method of the present application can be applied to training large models that perform complex reasoning tasks. According to the task scenario applicable to the large model to be trained, question-answering training data under the corresponding task scenario can be selected.
[0067] Taking the mathematics tutoring model as an example, mathematics problems can be used as problem samples, and the problem-solving process and answers of mathematics problems can be used as answer labels corresponding to the problem samples.
[0068] Step S110: Send the problem sample to the large model to be trained for inference to obtain the prediction output of the large model.
[0069] Specifically, the large model is fed with a sample question, which then performs the inference process and generates a prediction output. For complex reasoning tasks, the prediction output of the large model can include both the thought content and the answer.
[0070] Step S120: Calculate exploration rewards based on the hidden layer state features generated by the large model inference process. The exploration rewards are used to encourage the large model to adopt unknown inference paths to process problem samples.
[0071] Specifically, the large model inference process will generate corresponding hidden layer state features s at different times t Among them, s t Represents the hidden state features generated by the large model at time t. The hidden state features are used to guide the generation of the output character token at the corresponding time.
[0072] In this step, the exploration reward can be calculated based on the hidden layer state features generated in the large model inference process, and the exploration reward is used to encourage the large model to adopt an unknown inference path to process the input problem sample. Since the hidden layer state features can be at different times, the exploration rewards at different times can be calculated based on the hidden layer state features at different times, and the exploration rewards at different times can guide the large model to better adjust the inference path in the middle step of inference, that is, step-by-step path optimization can be achieved.
[0073] Step S130, calculating the result reward based on the predicted output and the answer label.
[0074] It should be noted that steps S120 and S130 can be executed in parallel, or in any order, Figure 2 Only an optional implementation process is shown.
[0075] In this step, the answer label corresponding to the problem sample is taken as the result supervision signal, and the result reward is calculated with the predicted output of the large model. The result reward is used to encourage the predicted output of the large model to approach the answer label.
[0076] Step S140, updating the parameters of the large model in the reinforcement learning manner according to the exploration reward and the result reward.
[0077] In this embodiment, the reward model includes both the exploration reward and the result reward, which can guide the adjustment of the inference path from the dimensions of the predicted output of the large model and the middle steps of the inference process of the large model, and update and train the parameters of the large model in the reinforcement learning manner.
[0078] For example, the value of the advantage function (defined as advantage) can be calculated according to the result reward and the exploration reward, and then the parameters of the large model can be updated according to the advantage.
[0079] For another example, according to the traditional reinforcement learning training method, the initial advantage A old can be obtained based on the result reward. Further, the exploration reward new can be injected into the advantage estimation to maintain the integrity of the result-driven advantage distribution. Then the new advantage A new is:
[0080] A old +
[0081] Wherein, represents the exploration reward of the large model inference process at the t-th time.
[0082] The above method ensures that the exploration reward does not interfere with the result-driven advantage estimation, thereby maintaining the integrity of the policy gradient.
[0083] The large model training method provided by the embodiments of the present application trains the large model by using the question and answer training data. On the basis of calculating the result reward based on the prediction output of the large model and the answer label, the exploration reward is further calculated based on the hidden layer state features generated in the inference process of the large model. The exploration reward is used to encourage the large model to process the input problem sample by using unknown inference paths. Meanwhile, the exploration reward and the result reward are referred to, and the parameters of the large model are updated in the reinforcement learning manner. As can be seen, the exploration reward is additionally added in the reinforcement learning process, which can encourage the large model to explore unknown inference paths, avoid falling into local optimum, and improve the probability of the large model finding a better inference path when facing a complex problem, thereby improving the performance of the large model on a complex inference task.
[0084] Further, since the exploration reward is calculated based on the hidden layer state features generated in the inference process of the large model, the exploration reward can be calculated based on the hidden layer state features generated at different time points in the inference process of the large model. The exploration rewards at different time points can guide the large model to better adjust the inference path in the middle step of inference, that is, the step-by-step path optimization can be achieved, and the inference performance of the large model is further improved, especially on a complex inference task.
[0085] In combination Figure 3 which illustrates an improved reinforcement learning training framework diagram.
[0086] Among them, the large model to be trained is taken as a policy network, and a prediction network is introduced. The exploration reward is realized by the uncertainty of the prediction network. The prediction network is configured to take the hidden layer state features s t as input and output the probability distribution of the prediction result. The exploration reward can be calculated based on the probability distribution. The prediction network can adopt a Bayesian network to output a probability distribution, such as a normal distribution, a Poisson distribution, etc.
[0087] In the reinforcement learning framework, the policy network receives the environment input to generate an action (prediction output). The reward model evaluates the action to obtain a result reward. The policy network updates the parameters through an update mechanism according to the reward (exploration reward and result reward) and generates a new action again, and the cycle continues until the training is completed. The update mechanism can adopt various reinforcement learning strategies, such as PPO algorithm, GRPO algorithm, etc.
[0088] In a possible implementation, the process of calculating the exploration reward based on the hidden layer state features generated in the inference process of the large model includes:
[0089] The hidden layer state features s t generated in the inference process of the large model are sent into the configured prediction network to obtain the probability distribution of the prediction result.
[0090] A first statistic is calculated based on the probability distribution, the first statistic representing uncertainty of the prediction network. That is, the prediction network predicts uncertainty of the result.
[0091] In an optional example, the first statistic can be represented by variance of the probability distribution. Variance reflects the dispersion of data, the greater the variance, the more dispersed the data, indicating that the prediction network has higher uncertainty of the prediction result.
[0092] In other optional examples, the first statistic can also be represented by other indicators capable of reflecting the dispersion of data, including but not limited to: standard deviation, interquartile range, range, mean absolute deviation, coefficient of variation, etc.
[0093] Based on the first statistic, an exploration reward is determined, the size of the exploration reward being in positive correlation with the first statistic.
[0094] In an optional example, the first statistic can be directly defined as the exploration reward. Alternatively, the first statistic can be weighted to obtain the exploration reward.
[0095] Taking the probability distribution output by the prediction network as a normal distribution as an example, it is represented as:
[0096] ( ) ~ .
[0097] Wherein, ( ) represents the probability distribution output by the prediction network when the hidden layer state feature s t is input, which conforms to the normal distribution.
[0098] is the mean, is the variance, representing the uncertainty of the prediction.
[0099] When the uncertainty of the prediction network is high, a higher exploration reward can be given to encourage the model to try new reasoning paths. The exploration reward can be represented as:
[0100] = .
[0101] Wherein, is a hyperparameter, used to control the weight of the exploration reward.
[0102] Further optionally, the process of calculating the exploration reward can also include updating the parameters of the prediction network.
[0103] The updating of the prediction network can be achieved by minimizing the prediction error and maximizing the exploration reward, where the prediction error is minimized to make the prediction result of the prediction network closer to the true target value (the answer label corresponding to the problem sample), i.e., the prediction network is trained to approximate the policy network, and the exploration reward is maximized to give a higher exploration reward when the uncertainty of the prediction network is high, thereby encouraging the model to try new inference paths.
[0104] Specifically, the process of updating the prediction network parameters can include:
[0105] S1, calculating a second statistical quantity based on the probability distribution output by the prediction network, the second statistical quantity representing the prediction result.
[0106] In an example, the second statistical quantity can be the mean of the probability distribution. Taking the above probability distribution as a normal distribution as an example, the second statistical quantity can be the mean . In addition, the second statistical quantity can also use other statistical indicators.
[0107] S2, calculating a prediction loss with the goal of making the second statistical quantity approach the answer label. The prediction loss can be represented as: .
[0108] where MSE represents the mean square error, used to calculate the mean square error between the mean t at the t-th time and the answer label y
[0109] S3, calculating a total loss according to the prediction loss and the exploration reward.
[0110] where the total loss is positively correlated with the prediction loss and negatively correlated with the exploration reward.
[0111] In an optional example, the total loss can be calculated according to the prediction loss, the exploration reward, and a weight factor λ, which is used to balance the weights of the prediction loss and the exploration reward in the total loss. For example, the total loss can be represented as: -λ .
[0112] where λ is used to balance the weights of the prediction loss and the exploration reward.
[0113] In some possible implementations, the weight factor λ can be set as a fixed parameter or an adjustable parameter. In the training process of a large model, a larger λ value can be set at the beginning of training to encourage the model to explore more; and λ can be gradually reduced at the later stage of training to make the model use existing knowledge for inference more, thereby ensuring the stability and adaptability of the training process.
[0114] Exemplarily, the large model training process includes at least two stages, and the weight of the exploration reward in the second stage is smaller than the weight of the exploration reward in the first stage by adjusting the size of the weight factor λ, and the second stage is later in time sequence than the first stage.
[0115] S4, updating the parameters of the prediction network according to the total loss.
[0116] The parameters of the prediction network are updated by aiming to minimize the prediction loss and maximize the exploration reward, so that the prediction network can approximate the policy network while encouraging the policy network to try new reasoning paths and avoid falling into a local optimal solution.
[0117] In some embodiments of the present application, before the aforementioned embodiment step S140, the parameters of the large model are updated in a reinforcement learning manner according to the exploration reward and the result reward, the following processing step can be added:
[0118] The current exploration reward calculated is decayed according to the current training step number to obtain a decayed exploration reward, and the decay degree of the exploration reward becomes larger with the increase of the training step number.
[0119] Define γ as the decay rate and n as the current training step number, then the decayed exploration reward can be expressed as:
[0120] .
[0121] As can be seen from the above formula, the decay degree of the exploration reward becomes larger with the increase of the training step number.
[0122] Therefore, step S140 is specifically updating the parameters of the large model in a reinforcement learning manner according to the decayed exploration reward and the result reward.
[0123] The method provided in the embodiment gradually reduces the intensity of the exploration reward as the large model training proceeds, so as to guide the large model to use the known correct reasoning path for reasoning more in the later training stage, rather than excessive exploration, thereby ensuring the stability and adaptability of the training.
[0124] When dealing with complex reasoning tasks, the current large model usually tends to generate longer reasoning chains (thinking content) in order to pursue higher accuracy, even if it is a simple problem, it will also develop a complex reasoning process, thereby damaging its problem solving efficiency. Therefore, the embodiment hopes to alleviate the phenomenon of excessive thinking of the large model without weakening the reasoning ability of the large model, and at the same time encourage the large model to explore unknown paths and improve the performance in complex reasoning tasks.
[0125] In the embodiment, the question and answer process is divided into three stages: fast thinking, verification, and slow thinking. This reduces unnecessary reflection of large models when dealing with simple tasks, while improving the reasoning ability of large models when dealing with complex tasks.
[0126] In this scheme, the fast thinking stage requires the model to quickly generate a preliminary answer within a strict character token budget, allowing it to quickly identify promising reasoning paths within limited resources. Not only does this improve reasoning speed, but it also reduces unnecessary computation on simple tasks. The verification stage evaluates the preliminary answer generated by the fast thinking stage to determine its correctness. This design not only provides immediate feedback to help large models quickly identify errors, but also avoids wasting excessive computing resources on incorrect paths. Through the verification stage, the large model can more effectively utilize computing resources, improving the accuracy and efficiency of reasoning. If the answer from the fast thinking stage is verified as incorrect, the large model enters the slow thinking stage to revise the preliminary answer, otherwise it directly outputs the correct preliminary answer. The slow thinking stage allows the large model to use more token budget for in-depth thinking, thereby improving the accuracy of the answer. In the slow thinking stage, the large model training method introduced in the foregoing embodiment can be used, that is, in addition to considering the result reward, an exploration reward is additionally considered, thereby encouraging the large model to explore unknown reasoning paths in the slow thinking stage, improving the probability of the large model discovering better reasoning paths when facing complex problems, and thereby improving the performance of the large model on complex reasoning tasks.
[0127] This phased reasoning approach in the embodiment avoids using high computational cost slow thinking throughout the entire reasoning process, thereby improving overall reasoning efficiency. The specific steps are shown in Figure 4 , which include:
[0128] Step S200, fast thinking stage: taking the question sample as the input of the large model, instructing the large model to quickly generate an initial answer, and calculating a first reward based on the initial answer and the target answer corresponding to the question sample.
[0129] The fast thinking stage can limit the token budget applicable to the generation of the initial answer by the large model, with the purpose of allowing the large model to quickly generate the initial answer within a limited token budget and avoid overthinking. The generated initial answer is usually a direct guess result or a heuristic-based preliminary solution for the question sample, etc.
[0130] The first reward of the fast thinking stage can be set as a binary reward function, that is, if the initial answer and the target answer are consistent, the reward is 1, otherwise the reward is 0.
[0131] After obtaining the first reward, the large model can be trained for parameter update according to the first reward.
[0132] Through the fast thinking stage, the intuition ability of the large model can be trained to quickly identify a promising reasoning path within limited resources.
[0133] Step S210, verification stage: instruct the large model to verify the correctness of the initial answer, and calculate the second reward according to the correctness of the verification result output by the large model. In the case where the verification result indicates that the initial answer is incorrect, step S220 is entered.
[0134] Specifically, the initial answer output by the large model in the fast thinking stage is taken as the input of the large model in the verification stage, the large model is instructed to verify the correctness of the initial answer, the second reward is calculated according to the correctness of the verification result output by the large model, and in the case where the verification result indicates that the initial answer is incorrect, the slow thinking stage is entered.
[0135] The goal of the verification stage is to evaluate the correctness of the initial answer, and the output verification result can be a binary verification result indicating whether the initial answer is correct. Since both the initial answer and the target answer can be obtained, that is, based on the consistency of the initial answer and the target answer, the standard answer of whether the initial answer is correct can be obtained. The correctness of the verification result output by the large model in the verification stage can be judged based on the standard answer.
[0136] The second reward in the verification stage can be set as a binary reward function, that is, if the verification result is correct, the reward is 1, otherwise the reward is 0.
[0137] After obtaining the second reward, the large model can be trained according to the second reward.
[0138] Through the verification stage, the verification ability of the large model can be trained to quickly judge the reliability of the initial answer and avoid unnecessary further reasoning.
[0139] Step S220, slow thinking stage: taking the problem sample and the initial answer as an input sequence to form the input of the large model, instructing the large model to reason and generate a revised answer, and calculating the exploration reward and the result reward according to the exploration reward and the result reward, and updating the parameters of the large model in the reinforcement learning manner.
[0140] Specifically, in the slow thinking stage, the exploration reward is calculated based on the hidden layer state features generated by the reasoning process of the large model, and the exploration reward is used to encourage the large model to use unknown reasoning paths to process the input sequence. The result reward is calculated based on the revised answer generated by the large model and the target answer. According to the exploration reward and the result reward, the parameters of the large model are updated in the reinforcement learning manner.
[0141] Among them, the related calculation process of the exploration reward can refer to the related introduction of the foregoing embodiments, which will not be repeated here.
[0142] In the case that the verification result of the verification phase indicates that the initial answer is wrong, the slow thinking phase is entered. In the slow thinking phase, the large model can use more token budget for in-depth thinking, correct errors, and generate corrected answers.
[0143] In the above training process, the slow thinking phase plays an important role. If the preliminary answer generated by the fast thinking phase is determined to be wrong by the verification phase, the slow thinking phase is the only opportunity for the model to correct the error. Through slow thinking, the large model reconsiders the problem, analyzes the reasons for the error, and tries to generate a more accurate answer. In addition, for complex reasoning tasks, the fast thinking phase may not be able to generate an accurate answer, while the slow thinking phase can use more token budget for in-depth reasoning to gradually approach the correct answer. The slow thinking phase not only needs to correct errors to provide learning opportunities for the large model, but also needs to analyze errors and generate correct answers so that the large model can better understand the structure of the problem and the problem-solving method, thereby performing better in future reasoning.
[0144] Existing reinforcement learning algorithms rely on sparse result-oriented rewards, which are difficult to provide effective feedback on complex problems, and the existing reward structure is not conducive to true exploration, and the model tends to use known high-reward paths rather than discover new reasoning paths. Therefore, the exploration reward-based reinforcement learning training strategy is introduced in the slow thinking phase to optimize the performance of the slow thinking process. The exploration reward-based reinforcement learning training strategy can refer to the related descriptions of the foregoing embodiments, which will not be described here.
[0145] The large model training method provided in this embodiment combines the phased reasoning method with the exploration reward-based reinforcement learning training strategy, avoiding the use of high computational cost slow thinking throughout the entire reasoning process through the phased reasoning method, thereby improving the overall reasoning efficiency. At the same time, through the exploration reward-based reinforcement learning training strategy, the large model can be encouraged to explore unknown reasoning paths in the slow thinking phase, improving the probability of the large model discovering better reasoning paths when facing complex problems, and thereby improving the performance of the large model on complex reasoning tasks.
[0146] Further, the large model training method provided in this embodiment enters the slow thinking phase only when the verification result output by the verification phase indicates that the preliminary answer is wrong, and the exploration reward is activated in this phase. That is, the exploration reward is activated only on the wrong samples to ensure that the model explores more on the wrong samples, improving the exploration efficiency.
[0147] In the case of the large model trained by the large model training method of any of the preceding embodiments, the present embodiment further provides a large model based question and answer method. The question and answer method of the present application can be applied to the system architecture as shown in Figure 1 , that is, the question and answer method can be executed by the terminal 100, or the question and answer method can be executed by the terminal 100 and the server 200 in cooperation.
[0148] The question and answer method of the present embodiment specifically includes the following steps:
[0149] First, obtain a question description.
[0150] Among them, according to different question and answer scenarios, the question description can be question description information corresponding to the question and answer scenario.
[0151] Further, the question description is sent to the configured large model to obtain the answer result output by the large model. The large model is trained by the large model training method of any of the preceding embodiments.
[0152] Based on the foregoing introduction of the large model training method, it can be known that the training method of the present application can improve the performance of the trained large model, especially the performance on complex reasoning tasks. Applying it to the question and answer task can improve the answer quality.
[0153] The large model training device provided by the present application embodiment is described below. The large model training device described below can be correspondingly referred to with the large model training method described above.
[0154] Referring to Figure 5 , Figure 5 , a large model training device structure diagram disclosed by the present application embodiment.
[0155] As shown in Figure 5 , the device can include:
[0156] The training data acquisition unit 11 is configured to acquire question and answer training data, wherein the question and answer training data includes question samples and corresponding answer labels;
[0157] The computing unit 12 is configured to send the question samples to a large model to be trained for reasoning to obtain a prediction output of the large model; calculate an exploration reward based on the hidden layer state features generated in the reasoning process of the large model, wherein the exploration reward is used to encourage the large model to process the question samples using unknown reasoning paths; calculate a result reward based on the prediction output and the answer labels; and update the parameters of the large model in a reinforcement learning manner according to the exploration reward and the result reward.
[0158] In an optional implementation, the process of calculating the exploration reward based on the hidden layer state features generated in the reasoning process of the large model includes:
[0159] The hidden state features generated by the large model inference process are fed into the configured prediction network to obtain the probability distribution of the prediction results;
[0160] calculating a first statistic based on the probability distribution, the first statistic representing uncertainty of the prediction network;
[0161] An exploration reward is determined based on the first statistic, and a size of the exploration reward is positively correlated with the first statistic.
[0162] In an optional implementation, the process of calculating the exploration reward based on the hidden state features generated by the large model inference process by the computing unit further includes:
[0163] calculating a second statistic based on the probability distribution, wherein the second statistic represents a prediction result;
[0164] Calculating a prediction loss with the goal of making the second statistic approach the answer label, and calculating a total loss based on the prediction loss and the exploration reward, where the total loss is positively correlated with the prediction loss and negatively correlated with the exploration reward;
[0165] The parameters of the prediction network are updated according to the total loss.
[0166] In an optional implementation, the process of calculating the total loss by the computing unit according to the prediction loss and the exploration reward includes:
[0167] Calculating a total loss based on the prediction loss, the exploration reward, and an adjustable weight factor, wherein the weight factor is used to balance the weights of the prediction loss and the exploration reward in the total loss;
[0168] The large model training process includes at least two stages. By adjusting the size of the weight factor, the weight of the exploration reward in the second stage is smaller than the weight of the exploration reward in the first stage. The second stage is later than the first stage in time sequence.
[0169] In an optional implementation, the first statistic is the variance of the probability distribution.
[0170] In an optional implementation, the second statistic is the mean of the probability distribution.
[0171] In an optional implementation, after calculating the exploration reward, the computing unit is further configured to: attenuate the exploration reward according to the current number of training steps to obtain an attenuated exploration reward, wherein the degree of attenuation of the exploration reward increases with an increase in the number of training steps; then, the process of updating the parameters of the large model using a reinforcement learning method according to the exploration reward and the result reward includes:
[0172] According to the decayed exploration reward and the result reward, the parameters of the large model are updated in a reinforcement learning manner.
[0173] Referring to Figure 6 , Figure 6 Another large model training device structure disclosed in the embodiments of the present application is shown in the figure.
[0174] As Figure 6 shown, the device can include:
[0175] The first computing unit 21 is configured to perform the training process in the fast thinking stage. Specifically, in the fast thinking stage, the large model is instructed to generate an initial answer based on the question sample as input, and a first reward is calculated based on the initial answer and the target answer corresponding to the question sample.
[0176] The second computing unit 22 is configured to perform the training process in the verification stage. Specifically, in the verification stage, the initial answer output by the large model in the fast thinking stage is used as input to the large model, and the large model is instructed to verify the correctness of the initial answer. A second reward is calculated according to the correctness of the verification result output by the large model, and in the case where the verification result indicates that the initial answer is incorrect, the slow thinking stage is entered.
[0177] The third computing unit 23 is configured to perform the training process in the slow thinking stage. Specifically, in the slow thinking stage, the question sample and the initial answer output by the large model in the fast thinking stage are combined to form an input sequence, which is used as input to the large model, and the large model is instructed to infer and generate a revised answer. An exploration reward is calculated based on the hidden layer state features generated during the inference process of the large model, and the exploration reward is used to encourage the large model to process the input sequence using unknown inference paths. A result reward is calculated based on the revised answer and the target answer. According to the exploration reward and the result reward, the parameters of the large model are updated in a reinforcement learning manner.
[0178] The above-mentioned large model training device can be implemented in whole or in part by software, hardware and their combinations. The above-mentioned units can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned units.
[0179] The embodiments of the present application also provide an electronic device. Referring to Figure 7 , which shows a structure diagram suitable for implementing the electronic device in the embodiments of the present application. The electronic device in the embodiments of the present application can include but is not limited to fixed terminals such as mobile phones, tablet computers, teaching large screens, wearable devices, etc. Figure 7The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0180] As shown in Figure 7 The electronic device can include a processing device (e.g., a central processor, a graphics processor, etc.) 1 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 2 or programs loaded from a storage device 8 into a random access memory (RAM) 3 to implement the large model training method or the question and answer method based on a large model of the aforementioned embodiments of the present application. In a state where the electronic device is powered on, various programs and data required for operation of the electronic device are also stored in the RAM 3. The processing device 1, the ROM 2, and the RAM 3 are connected to each other through a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.
[0181] Generally, the following devices can be connected to the I / O interface 5: input devices 6 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 7 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 8 including, for example, a memory card, a hard disk, etc.; and communication devices 9. The communication devices 9 can allow the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 The electronic device with various devices is shown, but it should be understood that it is not required to implement or have all the devices shown. More or fewer devices can be alternatively implemented or provided.
[0182] The embodiments of the present application also provide a computer program product including computer readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the large model training methods or the question and answer methods based on a large model provided by the embodiments of the present application.
[0183] The embodiments of the present application also provide a computer readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can cause the electronic device to implement any of the large model training methods or the question and answer methods based on a large model provided by the embodiments of the present application.
[0184] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate units can or can not be physically separate, and the units displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and the necessary general hardware, and of course can also be realized by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and the specific hardware structure for realizing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of software products, which are stored in readable storage media, such as computer floppy disks, U disks, mobile hard disks, ROM, RAM, magnetic or optical disks, etc., including a plurality of instructions for making a computer device (which can be a personal computer, a training device, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0186] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of a computer program product in whole or in part.
[0187] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0188] The various embodiments in the specification are described in a progressive manner, each embodiment focuses on the difference from other embodiments, the various embodiments can be combined as needed, and the same or similar parts refer to each other.
Claims
1. A large model training method, characterized in that: include: Obtaining question-answering training data, wherein the question-answering training data includes question samples and corresponding answer labels; Send the problem sample to the large model to be trained for inference to obtain the prediction output of the large model; Calculate exploration rewards based on the hidden state features generated by the large model inference process, and use the exploration rewards to encourage the large model to adopt unknown inference paths to process the problem samples; Calculating a result reward based on the predicted output and the answer label; According to the exploration reward and the result reward, the parameters of the large model are updated using reinforcement learning.
2. The method according to claim 1, characterized in that The process of calculating exploration rewards based on the hidden state features generated by the large model inference process includes: The hidden state features generated by the large model inference process are fed into the configured prediction network to obtain the probability distribution of the prediction results; calculating a first statistic based on the probability distribution, the first statistic representing uncertainty of the prediction network; An exploration reward is determined based on the first statistic, and a size of the exploration reward is positively correlated with the first statistic.
3. The method according to claim 2, characterized in that The process of calculating the exploration reward based on the hidden state features generated by the large model inference process also includes: calculating a second statistic based on the probability distribution, wherein the second statistic represents a prediction result; Calculating a prediction loss with the goal of making the second statistic approach the answer label, and calculating a total loss based on the prediction loss and the exploration reward, where the total loss is positively correlated with the prediction loss and negatively correlated with the exploration reward; The parameters of the prediction network are updated according to the total loss.
4. The method according to claim 3, characterized in that The process of calculating the total loss based on the prediction loss and the exploration reward includes: Calculating a total loss based on the prediction loss, the exploration reward, and an adjustable weight factor, wherein the weight factor is used to balance the weights of the prediction loss and the exploration reward in the total loss; The large model training process includes at least two stages. By adjusting the size of the weight factor, the weight of the exploration reward in the second stage is smaller than the weight of the exploration reward in the first stage. The second stage is later than the first stage in time sequence.
5. The method according to claim 3, characterized in that The first statistic is the variance of the probability distribution, and the second statistic is the mean of the probability distribution.
6. The method according to any one of claims 1 to 5, characterized in that After calculating the exploration reward, it also includes: Attenuating the exploration reward according to the current number of training steps to obtain an attenuated exploration reward, wherein the degree of attenuation of the exploration reward increases with the increase in the number of training steps; Then, according to the exploration reward and the result reward, the process of updating the parameters of the large model using reinforcement learning includes: According to the attenuated exploration reward and the result reward, the parameters of the large model are updated using reinforcement learning.
7. A large model training method, characterized in that: include: The training process consists of three stages: fast thinking stage, verification stage and slow thinking stage, among which, In the fast thinking stage: the question sample is used as the input of the large model, and the large model is instructed to quickly generate an initial answer. The first reward is calculated based on the initial answer and the target answer corresponding to the question sample; In the verification phase: the initial answer output by the large model in the fast thinking phase is used as input to the large model, and the large model is instructed to verify the correctness of the initial answer. The second reward is calculated based on the correctness of the verification result output by the large model. If the verification result shows that the initial answer is incorrect, the slow thinking phase is entered; During the slow thinking phase: The question sample and the initial answer output by the large model in the fast thinking stage form an input sequence as the input of the large model, and instruct the large model to infer and generate a revised answer; Calculating an exploration reward based on the hidden layer state features generated during the large model inference process, wherein the exploration reward is used to encourage the large model to adopt an unknown inference path to process the input sequence; Calculating a result reward based on the corrected answer and the target answer; According to the exploration reward and the result reward, the parameters of the large model are updated using reinforcement learning.
8. A question-answering method based on a large model, characterized in that: include: Get a description of the problem; The problem description is sent to the configured big model to obtain the answer result output by the big model, wherein the big model is trained using the big model training method of any one of claims 1-7.
9. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the large model training method as described in any one of claims 1 to 7, or to implement the various steps of the large model-based question-answering method as described in claim 8.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the various steps of the large model training method as described in any one of claims 1 to 7, or implements the various steps of the large model-based question-answering method as described in claim 8.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the various steps of the large model training method as described in any one of claims 1 to 7, or implements the various steps of the large model-based question-answering method as described in claim 8.
Citation Information
Cited By
Organ transplantation clinical aid decision-making method, device and equipment based on AI large model, medium and product
CN121096601A
Multi-modal large model training method and related device
CN121257750A
Structural semantic flow modeling-based interpretable text question and answer method and system
CN121365742A