Large Language Model Training Method, Logical Problem Handling Method and Computer Device
By generating thinking chain data and adjusting model parameters, the logical reasoning and generalization capabilities of large language models are improved, and the problem of poor handling of complex logical reasoning problems in the existing technology is solved.
Patent Information
- Application Number
- CN202510163560.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-14
AI Technical Summary
Existing large language models have poor results in dealing with complex logical reasoning problems and need to rely on techniques such as prompt engineering or multi-scheme search.
By obtaining the large language model with output formatting, thinking chain data is generated for logical problems concentrated in logic problems, and model parameters are adjusted based on this data to obtain a model with preliminary logical reasoning ability. Then, using a richer logical problem dataset and inference step scoring model, further adjust the model parameters and enhance its generalization ability.
The accuracy and generalization ability of large language models in handling complex logical problems is improved, so that they can perform logical reasoning and deal with unseen logical problems more effectively.
Smart Images

Figure CN119622344B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to large language model training methods, logical problem processing methods, and computer devices. Background Art
[0002] Large Language Models (LLMs) exhibit two thinking modes when processing information: fast thinking and slow thinking. Fast thinking relies on intuition and pattern recognition and is suitable for scenarios that require quick responses; while slow thinking involves complex logical reasoning and is suitable for problems that require in-depth analysis and reasoning.
[0003] Currently, LLMs have made significant progress in fast thinking. However, for solving complex logical reasoning problems, LLMs in related technologies often need to rely on techniques such as prompt engineering or multi-scheme search, and the processing effects of these techniques in solving complex logical reasoning problems are poor. Summary of the Invention
[0004] In view of this, the present invention provides a large language model training method, a logical problem processing method, and a computer device to solve the problem that the large language model in related technologies has a poor processing effect on complex logical problems.
[0005] In a first aspect, the present invention provides a large language model training method, and the method includes:
[0006] Obtain a first large language model with formatted output;
[0007] Obtain a first set of logical problems and multiple sets of hyperparameters of the first large language model;
[0008] For any first logical problem in the first set of logical problems, input the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem. The thought chain data includes multiple reasoning steps, and the large language model generates the current reasoning step based on the first logical problem, the previous reasoning step, and the evaluation information of the previous reasoning step;
[0009] Based on the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data, adjust the model parameters of the first large language model to obtain a second large language model with preliminary logical reasoning ability and a reasoning step scoring model;
[0010] Obtain a second set of logical problems, and based on the second set of logical problems and the reasoning step scoring model, adjust the model parameters of the second large language model to obtain a target large language model with generalization ability, where the amount of logical problem data in the second set of logical problems is more than that in the first set of logical problems.
[0011] For the large language model training method provided in this embodiment, for any first logical problem in the first logical problem set, the first logical problem is input into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem. The thought chain data includes multiple reasoning steps. The large language model generates the current reasoning step based on the first logical problem, the previous reasoning step, and the evaluation information of the previous reasoning step, ensuring that the generated thought chain data has actions of making mistakes and correcting mistakes. Furthermore, based on the first logical problem, the first logical chain data, and the evaluation information of the reasoning steps in the thought chain data, the model parameters of the first large language model are adjusted, and the obtained second large language model has more accurate and reasonable logical reasoning capabilities. Obtain a second logical problem set with a larger amount of logical problem data, and based on the second logical problem set and the reasoning step scoring model, adjust the model parameters of the second large language model, so that the second large language model can be exposed to more diverse logical problem scenarios and learn more extensive logical patterns and rules, thereby enhancing the generalization ability of the model when facing new and unseen logical problems. The target large language model obtained by using the large language model training method of this embodiment can accurately process complex logical problems.
[0012] In an alternative embodiment, obtaining the first large language model with formatted output includes:
[0013] Obtain a data set, where the training samples in the data set include logical problems and the reasoning steps corresponding to the logical problems;
[0014] Based on the data set, adjust the model parameters of the pre-trained large language model to obtain the first large language model with formatted output.
[0015] For the large language model training method provided in this embodiment, by adjusting the model parameters of the pre-trained large language model to obtain the first large language model with formatted output, the interpretability and transparency of the model output are improved, the format consistency and standardization of the model output results are ensured, and misunderstandings or errors caused by inconsistent output results are reduced.
[0016] In an alternative embodiment, based on the data set, adjusting the model parameters of the pre-trained large language model to obtain the first large language model with formatted output includes:
[0017] Input the training samples into the pre-trained large language model to obtain a first output result, where the first output result represents the probability value of the label marked by the pre-trained large language model as corresponding to the training sample;
[0018] Based on the first output result, determine a first loss value, where the first loss value includes the loss value of format symbols;
[0019] Based on the first loss value, adjust the model parameters of the pre-trained large language model until the adjustment termination condition is reached, and obtain the first large language model with formatted output.
[0020] The large language model training method provided in this embodiment obtains the loss value of the format symbol, and adjusts the model parameters of the pre-trained large language model according to the loss value of the format symbol, so as to obtain the first large language model with formatted output, improving the interpretability and transparency of the model output, ensuring the format consistency and standardization of the model output results, and reducing misunderstandings or errors caused by inconsistent output formats.
[0021] In an alternative embodiment, determining the first loss value based on the first output result includes:
[0022] Determine the position marked by the format symbol based on the reasoning steps corresponding to the logical problems in the training samples;
[0023] Based on the position marked by the format symbol, determine the probability value that the mark corresponding to the position marked by the format symbol in the first output result is the mark corresponding to the format symbol;
[0024] Determine the first loss value based on the probability value that the mark corresponding to the position marked by the format symbol is the mark corresponding to the format symbol.
[0025] The large language model training method provided in this embodiment determines the first loss value through the probability value that the mark corresponding to the position marked by the format symbol determined in the first output result is the mark corresponding to the format symbol, adjusts the model parameters of the pre-trained large language model based on the first loss value, and obtains the first large language model with formatted output, ensuring the formatting and consistency of the model output.
[0026] In an alternative embodiment, inputting the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem includes:
[0027] Input the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the first reasoning steps;
[0028] Input the first logical problem and the multiple first reasoning steps corresponding to the first logical problem into the third large language model to obtain the target first reasoning steps and the evaluation information of the target first reasoning steps, where the logical reasoning ability of the third large language model is higher than that of the pre-trained large language model;
[0029] Determine the target hyperparameters corresponding to the target first reasoning steps based on the target first reasoning steps;
[0030] Input the first logical problem, the target first reasoning step, and the evaluation information of the target first reasoning step into the first large language model corresponding to the target hyperparameters to obtain the second reasoning step;
[0031] Use the third large language model to determine the evaluation information of the second reasoning step based on the target first reasoning step, the second reasoning step, and the first logical problem;
[0032] Input the target first reasoning step, the second reasoning step, the first logical problem, and the evaluation information of the second reasoning step into the first large language model corresponding to the target hyperparameters to obtain the third reasoning step;
[0033] Iteratively execute using the third large language model to determine the evaluation information of the currently obtained reasoning step based on the first logical problem, the reasoning steps before the currently obtained reasoning step, and the currently obtained reasoning step, and input the first logical problem, the reasoning steps before the currently obtained reasoning step, the currently obtained reasoning step, and the evaluation information of the currently obtained reasoning step into the first language model corresponding to the target hyperparameters to obtain the next reasoning step until the logical reasoning ends to obtain the thought chain data of the first logical problem;
[0034] Among them, the evaluation information includes a score and guiding opinions.
[0035] The large language model training method provided in this embodiment inputs the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the first reasoning step, and then uses the third large language model with stronger logical reasoning ability to screen the first reasoning step to obtain the target first reasoning step and the evaluation information of the target first reasoning step, reducing the calculation amount of generating thought chain data while ensuring the diversity of the generated thought chain data.
[0036] In each iteration process, based on the currently obtained reasoning step and its evaluation information, the subsequent reasoning steps are further adjusted and optimized to ensure the logical consistency, coherence, and accuracy of the finally output thought chain data, ensuring that the generated thought chain data has actions of making mistakes and correcting mistakes, thereby improving the logical reasoning ability of the second large language model and improving the processing effect of complex logical problems.
[0037] In an optional implementation manner, inputting the first logical problem and multiple first reasoning steps corresponding to the first logical problem into the third large language model to obtain the target first reasoning step includes:
[0038] Use the third large language model to screen the multiple first reasoning steps corresponding to the first logical problem from the dimensions of correctness, uncertainty, and repeatability based on the first logical problem to obtain the target first reasoning step.
[0039] The large language model training method provided in this embodiment uses a third large language model to screen multiple first reasoning steps corresponding to the first logical problem from the dimensions of correctness, uncertainty, and repeatability, obtain the target first reasoning steps, remove the first reasoning steps that do not meet the requirements, and ensure the accuracy of the obtained thought chain data.
[0040] In an alternative embodiment, based on the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data, the model parameters of the first large language model are adjusted to obtain a second large language model with preliminary logical reasoning ability and a reasoning step scoring model, including:
[0041] Input the first logical problem and the thought chain data of the first logical problem into the first large language model to obtain a second output result, where the second output result represents the probability value of the labels output by the first large language model marked as the labels corresponding to the first logical problem and the thought chain data of the first logical problem;
[0042] Based on the second output result, determine the second loss value, where the second loss value includes the loss value of the reasoning steps of the thought chain data;
[0043] Based on the second loss value, adjust the model parameters of the first large language model to obtain a second large language model with preliminary logical reasoning ability;
[0044] Input the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data into the first large language model to obtain a third output result, where the third output result represents the probability value of the labels output by the first large language model marked as the labels corresponding to the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data;
[0045] Based on the third output result, determine the third loss value, where the third loss value includes the loss value of the score of the reasoning steps;
[0046] Based on the third loss value, adjust the model parameters of the first large language model to obtain a reasoning step scoring model;
[0047] Among them, the evaluation information includes scores.
[0048] The large language model training method provided in this embodiment calculates the loss value of the reasoning steps of the thought chain data and the loss value of the score of the reasoning steps, and adjusts the model parameters of the first large language model respectively based on the loss value of the reasoning steps of the thought chain data and the loss value of the score of the reasoning steps, improving the logical reasoning ability of the model and the accuracy of the reasoning steps.
[0049] In an alternative embodiment, based on the second logical problem set and the inference step scoring model, the model parameters of the second large language model are adjusted to obtain a target large language model with generalization ability, including:
[0050] Create multiple environments that share the second large language model;
[0051] In each environment, for any second logical problem in the second logical problem set, input the second logical problem into the second large language model to obtain the first inference step of the second logical problem;
[0052] Obtain the probability value of the first inference step;
[0053] Input the first inference step and the second logical problem into the inference step scoring model to obtain the score of the first inference step;
[0054] Input the second logical problem and the first inference step into the second large language model to obtain the second inference step of the second logical problem;
[0055] Obtain the probability value of the second inference step;
[0056] Input the second inference step, the second logical problem, and the first inference step into the inference step scoring model to obtain the score of the second inference step;
[0057] Iteratively execute inputting the second logical problem and the inference steps before the inference step to be obtained into the second large language model to obtain the inference step to be obtained, obtain the probability value of the inference step to be obtained, input the second logical problem, the inference step to be obtained, and the inference steps before the inference step to be obtained into the inference step scoring model to obtain the score of the inference step to be obtained until the logical reasoning of the environment for the second logical problem ends;
[0058] Save the second logical problem, the inference steps of the second logical problem, the probability values of the inference steps, and the scores of the inference steps during the logical reasoning process;
[0059] Based on the saved second logical problem, the inference steps of the second logical problem, the probability values of the inference steps, and the scores of the inference steps, adjust the model parameters of the second large language model to obtain a target large language model with generalization ability.
[0060] The large language model training method provided in this embodiment creates multiple environments, enabling the model to be fully trained in different environments, enhancing the generalization ability of the model, enabling the model to provide reliable and accurate answers when facing unseen logical problems, and improving its applicability in various scenarios.
[0061] In an alternative embodiment, obtaining the probability value of the first inference step includes:
[0062] Input the second logical problem and the first inference step into the second large language model to obtain the probability value of the first inference step.
[0063] The large language model training method provided in this embodiment enhances the logical reasoning ability of the model by obtaining the probability value of the first inference step and guiding the optimization of the model based on the probability value of the first inference step.
[0064] In an alternative embodiment, based on the saved second logical problem, the inference steps of the second logical problem, the probability values of the inference steps, and the scores of the inference steps, adjusting the model parameters of the second large language model to obtain a target large language model with generalization ability, including:
[0065] Based on the saved second logical problem, the inference steps of the second logical problem, and the probability values of the inference steps, determine the adjacent inference steps of the second logical problem under the reference model and the probability value of each inference step in the adjacent inference steps, where the reference model is the second large language model without adjusted model parameters;
[0066] Based on the saved scores of the inference steps, determine the score of each inference step in the adjacent inference steps;
[0067] Based on the adjacent inference steps of the second logical problem under the reference model, the probability value of each inference step in the adjacent inference steps, and the score of each inference step in the adjacent inference steps, determine the fourth loss value;
[0068] Based on the fourth loss value, adjust the model parameters of the second large language model to obtain a target large language model with generalization ability.
[0069] The large language model training method provided in this embodiment enhances the logical reasoning ability and generalization ability of the model by determining the fourth loss value based on the adjacent inference steps of the second logical problem under the reference model, the probability value of each inference step in the adjacent inference steps, and the score of each inference step in the adjacent inference steps, and adjusting the model parameters of the second large language model based on the fourth loss value to obtain a target large language model with generalization ability.
[0070] In an alternative embodiment, determining the fourth loss value based on the adjacent inference steps of the second logical problem under the reference model, the probability value of each inference step in the adjacent inference steps, and the score of each inference step in the adjacent inference steps includes:
[0071] Obtain the probability value of each inference step in the adjacent inference steps of the second logical problem under the second large language model corresponding to the current model parameters;
[0072] For any one of the adjacent inference steps, based on the probability value of the inference step in the adjacent inference steps of the second large language model corresponding to the current model parameters for the second logical problem and the probability value of the inference step in the adjacent inference steps of the second logical problem in the reference model, determine the divergence between the reference model and the second large language model corresponding to the current model parameters under the inference step, so as to obtain the divergence between the reference model and the second large language model corresponding to the current model parameters under each inference step in the adjacent inference steps;
[0073] Based on the score of each inference step in the adjacent inference steps and the divergence between the reference model and the second large language model corresponding to the current model parameters under each inference step in the adjacent inference steps, determine the fourth loss value.
[0074] The large language model training method provided in this embodiment determines the fourth loss value through the score of each inference step in the adjacent inference steps and the divergence between the reference model and the second large language model corresponding to the current model parameters under each inference step in the adjacent inference steps. Based on the fourth loss value, the model parameters of the second large language model are adjusted to obtain a target large language model with generalization ability, which helps to guide the model to learn a more reasonable and accurate inference process, thereby improving the performance of the model in logical reasoning tasks and making the output inference steps more logical and in line with the actual situation.
[0075] In an alternative implementation, determining the fourth loss value based on the score of each inference step in the adjacent inference steps and the divergence between the reference model and the second large language model corresponding to the current model parameters under each inference step in the adjacent inference steps includes:
[0076] Based on the score of the first inference step in the adjacent inference steps and the divergence between the reference model and the second large language model corresponding to the current model parameters under the first inference step, determine the first reward value;
[0077] Based on the scores of each inference step in the adjacent inference steps, determine the second reward value of the second inference step in the adjacent inference steps;
[0078] Based on the score of the second inference step in the adjacent inference steps, the divergence between the reference model and the second large language model corresponding to the current model parameters under the second inference step, and the second reward value of the second inference step in the adjacent inference steps, determine the third reward value;
[0079] Based on the first reward value and the third reward value, determine the fourth loss value;
[0080] Wherein, the adjacent inference steps include a first inference step and a second inference step.
[0081] The large language model training method provided in this embodiment rewards the self-improving behavior that occurs during the inference process of the model, motivates the model to continuously optimize in subsequent inferences, and attempts to generate better inference steps, thereby gradually improving the overall performance and inference quality of the model.
[0082] In a second aspect, the present invention provides a method for processing logical problems, including:
[0083] Obtain the logical problem to be processed;
[0084] Input the logical problem to be processed into the target large language model trained by using the large language model training method of the first aspect or any corresponding implementation manner thereof, and obtain the thought chain data corresponding to the logical problem to be processed.
[0085] The logical problem processing method provided in this embodiment processes the logical problem to be processed by using the target large language model trained by using the large language model training method of the first aspect or any corresponding implementation manner thereof, achieving efficient and accurate processing of logical problems.
[0086] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the large language model training method of the first aspect or any corresponding implementation manner thereof or execute the logical problem processing method of the second aspect.
[0087] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the large language model training method of the first aspect or any corresponding implementation manner thereof or execute the logical problem processing method of the second aspect.
[0088] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the large language model training method of the first aspect or any corresponding implementation manner thereof or execute the logical problem processing method of the second aspect. Description of the Drawings
[0089] In order to more clearly illustrate the specific implementation manners of the present invention or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the specific implementation manners or the related art. Obviously, the following drawings are some implementation manners of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0090] Figure 1 It is a schematic flowchart of a large language model training method according to an embodiment of the present invention;
[0091] Figure 2 It is a schematic flowchart of another large language model training method according to an embodiment of the present invention;
[0092] Figure 3 It is a schematic flowchart of yet another large language model training method according to an embodiment of the present invention;
[0093] Figure 4 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0094] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0095] A large language model is a complex model system constructed by leveraging deep learning technologies and relying on massive data for training. Such models possess powerful capabilities, enabling them to accurately understand natural language and simultaneously generate smooth and natural language content.
[0096] When processing information, large language models exhibit two thinking modes, similar to fast thinking and slow thinking in human thinking. In the fast thinking mode, large language models rely on intuition and existing pattern recognition capabilities to quickly provide answers. This mode is suitable for scenarios that require quick responses, such as daily conversations and simple Q&A scenarios that do not require complex logical reasoning.
[0097] Slow thinking involves complex logical reasoning and analysis. In the slow thinking mode, large language models need to analyze problems, construct complex thinking processes, and gradually derive answers. This mode is suitable for complex problems that require in-depth analysis and reasoning, such as scientific problem-solving and long-form writing.
[0098] Large language models in related technologies have made significant progress in fast thinking and have good application effects in scenarios that require quick responses. However, for the processing of complex logical reasoning problems, that is, in slow thinking, large language models in related technologies often need to rely on techniques such as prompt engineering or multi-scheme search. These techniques have poor processing effects when solving complex logical reasoning problems such as those in science and mathematics.
[0099] Specifically, in the related art, large language models improve the logical reasoning ability of large language models through techniques such as Chain-of-Thought (COT), Tree of Thoughts (TOT), search inference enhancement technology, STaR method, Quiet-STaR method, etc.
[0100] Among them, COT is a technique to enhance the reasoning ability of LLM, aiming to improve the performance of the model in complex tasks by generating intermediate reasoning steps. The characteristics of COT include: Step-by-step reasoning: COT requires the model to gradually derive a series of intermediate steps before generating the final answer. These steps form a chain of thought, helping the model to better understand the problem and draw the correct conclusion. Interpretability: By presenting the reasoning results, COT improves the transparency of the model's decision-making, enabling users to understand how the model arrives at a certain conclusion. Application scope: The COT technique is applicable to various reasoning tasks, including arithmetic reasoning, common sense reasoning, and symbolic reasoning, etc. It helps the model to perform more effective logical reasoning when facing problems that require combining multiple pieces of information.
[0101] TOT aims to enhance the performance of LLM in complex reasoning tasks and is an evolution of the COT method. By introducing a tree structure to organize the logical thinking process, the model can consider multiple paths and options when solving problems, improving the model's reasoning ability and output quality.
[0102] The search inference enhancement technology can use the LLM to generate multiple answers during reasoning, and then determine the final answer by means of voting or in combination with a reward model.
[0103] The STaR method gradually guides the model to improve its complex reasoning ability by iteratively using a small number of reasoning examples and a large amount of non-reasoning data sets. This method relies on a simple loop: generate the reasoning processes of multiple questions and prompt with a small number of reasoning examples; if the generated answer is incorrect, then try to generate the reasoning process again with the correct answer given, and fine-tune all the reasoning processes that finally produce the correct answer, and repeat the above process.
[0104] The Quiet-STaR method allows the LLM to perform a series of internal thinking steps before generating the final answer, that is, to perform sufficient internal reasoning before the model outputs. These thinking steps can help the model to understand the problem more deeply, so as to generate a more accurate and comprehensive answer.
[0105] However, COT, TOT, and search inference enhancement techniques do not involve weight updates to the LLM, which means they rely on the LLM having long-range thinking and self-correction abilities during the pre-training and fine-tuning phases. However, this is clearly unrealistic. Therefore, the model performs poorly when dealing with complex logical problems. Long-range thinking ability refers to the ability to perform complex, coherent, and long-chain thinking and reasoning. Even if the training dataset for the LLM contains some inference chain data, the diversity and generalization ability of the obtained LLM are still poor, and the performance when dealing with complex logical problems is poor. STaR and Quiet-STaR techniques attempt to enhance the complex logical reasoning ability of the model through fine-tuning. Fine-tuning is to further train the model using a specific dataset based on the pre-trained model to optimize the performance of the model on specific tasks. However, the method of obtaining the specific dataset in these two ways is directly guided by COT. The long-range thinking ability and self-correction ability of the dataset obtained in this way only stay at the solution level rather than the step level, resulting in poor performance of the model when dealing with complex logical problems.
[0106] An embodiment of the present invention provides a large language model training method. For any first logical problem in the first logical problem set, the first logical problem is input into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem. The thought chain data includes multiple reasoning steps. The large language model generates the current reasoning step based on the first logical problem, the previous reasoning step, and the evaluation information of the previous reasoning step; based on the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data, the model parameters of the first large language model are adjusted to obtain a second large language model with preliminary logical reasoning ability and a reasoning step scoring model; a second logical problem set is obtained, and based on the second logical problem set and the reasoning step scoring model, the model parameters of the second large language model are adjusted to obtain a target large language model with generalization ability. Among them, the amount of logical problem data in the second logical problem set is more than that in the first logical problem set, achieving the effect of improving the logical reasoning ability and generalization ability of the large language model, so that the target large language model can accurately process complex logical problems.
[0107] According to an embodiment of the present invention, an embodiment of a large language model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0108] In this embodiment, a large language model training method is provided, which can be used in mobile terminals such as servers and central processing units. Figure 1is a flowchart of a large language model training method according to an embodiment of the present invention, as Figure 1 shown, the process includes the following steps:
[0109] Step S101, obtain a first large language model with output formatting.
[0110] Among them, obtain a pre-trained large language model in the related art, and obtain a first large language model with output formatting by fine-tuning the model parameters in the pre-trained large language model.
[0111] The model parameters include weight parameters.
[0112] Step S102, obtain a first set of logical problems and multiple sets of hyperparameters of the first large language model.
[0113] Among them, the first set of logical problems includes multiple first logical problems. A logical problem is a problem involving logical reasoning, and these problems usually require drawing conclusions through reasonable reasoning steps based on given preconditions, rules, or information.
[0114] Exemplarily, the logical problem can be: "A circle with a radius of 5 is externally tangent to a circle with a radius of 3. Connect the centers of the two circles and draw the perpendicular bisector of this line segment. If the intersection points of the perpendicular bisector and the two circles are A and C respectively, find the length of line segment AC".
[0115] The hyperparameters can be adjusted manually, and the hyperparameters can be the temperature, sampling range, etc. output by the first language model.
[0116] Exemplarily, the number of sets of hyperparameters can be 32 sets. Of course, it can also be other numbers of sets, and no specific limitation is made here.
[0117] Step S103, for any first logical problem in the first set of logical problems, input the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem. The thought chain data includes multiple reasoning steps, and the large language model generates the current reasoning step based on the first logical problem, the previous reasoning step, and the evaluation information of the previous reasoning step.
[0118] It can be understood that for any first logical problem, multiple thought chain data can be obtained.
[0119] When generating any reasoning step in the thought chain data, the large language model generates the reasoning step based on the first logical problem, the previous reasoning step (i.e., the reasoning step before this reasoning step) and the evaluation information of the previous reasoning step.
[0120] It should be noted that the evaluation information of the previous reasoning step includes the evaluation information of the most recent reasoning step in the reasoning steps before this reasoning step.
[0121] That is to say, when generating any reasoning step in the thought chain data, the large language model is based on the first logical problem, the reasoning steps before this reasoning step, i.e., the previous reasoning steps, and the evaluation information of the nearest reasoning step among the reasoning steps before this reasoning step, i.e., the evaluation information of the nearest reasoning step in the previous reasoning steps, to generate this reasoning step. Among them, the evaluation information includes a score and guiding opinions.
[0122] Step S104: Based on the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data, adjust the model parameters of the first large language model to obtain a second large language model with preliminary logical reasoning ability and a reasoning step scoring model.
[0123] Among them, the loss function used when obtaining the second large language model with preliminary logical reasoning ability is different from the loss function used when obtaining the reasoning step scoring model.
[0124] The loss function used when obtaining the second large language model with preliminary logical reasoning ability calculates the loss value of the reasoning step content output by the first large language model.
[0125] The loss function used when obtaining the reasoning step scoring model calculates the loss value of the score after each reasoning step.
[0126] Step S105: Obtain a second set of logical problems. Based on the second set of logical problems and the reasoning step scoring model, adjust the model parameters of the second large language model to obtain a target large language model with generalization ability. Among them, the amount of logical problem data in the second set of logical problems is more than that in the first set of logical problems.
[0127] Among them, the logical problems in the second set of logical problems can include some of the logical problems in the first set of logical problems or all of the logical problems in the first set of logical problems.
[0128] It should be noted that the second large language model has preliminary logical reasoning ability, that is, it has preliminary long thought chain ability, but its generalization ability is not strong enough. Therefore, it is necessary to use the reinforcement learning method to adjust the model parameters of the second large language model based on the second set of logical problems and the reasoning step scoring model to obtain a target large language model with generalization ability.
[0129] The large language model training method provided in this embodiment, for any first logical problem in the first logical problem set, inputs the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem. The thought chain data includes multiple reasoning steps. The large language model generates the current reasoning step based on the first logical problem, the previous reasoning steps, and the evaluation information of the previous reasoning steps, ensuring that the generated thought chain data has actions of making mistakes and correcting mistakes. Furthermore, based on the evaluation information of the reasoning steps in the first logical problem, the first logical chain data, and the thought chain data, the model parameters of the first large language model are adjusted, and the obtained second large language model has more accurate and reasonable logical reasoning capabilities. Obtain a second logical problem set with a larger amount of logical problem data, and based on the second logical problem set and the reasoning step scoring model, adjust the model parameters of the second large language model, enabling the second large language model to be exposed to more diverse logical problem scenarios and learn more extensive logical patterns and rules, thereby enhancing the generalization ability of the model when facing new and unseen logical problems. The target large language model obtained by using the large language model training method of this embodiment can accurately process complex logical problems.
[0130] In this embodiment, a large language model training method is provided, which can be used in mobile terminals, such as servers, central processing units, etc. Figure 2 It is a flowchart of the large language model training method according to an embodiment of the present invention, as Figure 2 shown, and this process includes the following steps:
[0131] Step S201, obtain a first large language model with formatted output.
[0132] Specifically, the above step S201 includes:
[0133] Step S2011, obtain a data set, and the training samples in the data set include logical problems and the reasoning steps corresponding to the logical problems.
[0134] Among them, the reasoning steps corresponding to the logical problems are that the logical problems are decomposed into a series of logically related steps. These reasoning steps can be generated using the way of COT and few-shot prompting engineering. This data set is used to fine-tune the pre-trained large language model.
[0135] Exemplarily, a training sample can be as follows:
[0136] Question: A circle with a radius of 5 is externally tangent to a circle with a radius of 3. Connect the centers of the two circles and draw the perpendicular bisector of this line segment. If the intersection points of the perpendicular bisector and the two circles are A and C respectively, find the length of the line segment AC.
[0137] The first step: Analyze the geometric relationship between the two circles and the distance between the centers.
[0138] According to the problem statement, two circles are externally tangent, and their radii are 5 and 3 respectively. Therefore, the distance between their centers is equal to the sum of the radii of the two circles, that is, 5 + 3 = 8. Denote the distance between the centers as O1O2 = 8.
[0139] (This " " is a symbol in terms of format)
[0140] Step 2: Find the geometric characteristics of the length of AC.
[0141] After connecting the centers O1 and O2 and drawing the perpendicular bisector of this line segment, A and C are respectively the intersection points of the two circles on the perpendicular bisector. This indicates that the line segment AC is perpendicular to O1O2 and lies on the perpendicular bisector of the line connecting the centers. Since A and C are respectively located on the large circle and the small circle, the distance from A to the center O1 is 5, and the distance from C to the center O2 is 3.
[0142]
[0143] Step 3: Determine the expression for the length of AC.
[0144] In the right triangle O1AO2, the sum of the distance from A to the line O1O2 and the distance from C to the line O1O2 is the length of AC, that is, AC = O1A + O2C = 5 + 3.
[0145]
[0146] Step 4: Calculate the length of AC.
[0147] Therefore, the length of AC is 5 + 3 = 8.
[0148]
[0149] Answer: The length of the line segment AC is 8.
[0150] It can be understood that the questions in the training samples are logical problems, and the logical steps corresponding to the logical problems in the training samples can be Step 1, Step 2, Step 3, Step 4, Answer, etc.
[0151] In step S2012, based on the data set, the model parameters of the pre-trained large language model are adjusted to obtain the first large language model with formatted output.
[0152] Among them, after obtaining the data set, the model parameters of the pre-trained large language model are adjusted using the data set to obtain the first large language model with formatted output. The pre-trained large language model is a large language model with natural language text understanding and processing capabilities in related technologies.
[0153] The first large language model with output formatting is a large language model capable of generating formatted reasoning steps. The formatted reasoning steps can refer to outputting a procedural solution line by line according to the reasoning steps for solving the problem until the problem is processed.
[0154] That is to say, the first large language model can output data according to the formatted reasoning steps.
[0155] Step S202, obtain a first set of logical problems and multiple sets of hyperparameters of the first large language model. For details, please refer to Figure 1 Step S102 of the illustrated embodiment, which will not be elaborated here.
[0156] Step S203, for any first logical problem in the first set of logical problems, input the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical problem. The thought chain data includes multiple reasoning steps, and the large language model generates the current reasoning step based on the first logical problem, the previous reasoning steps, and the evaluation information of the previous reasoning steps.
[0157] Specifically, the above step S203 includes:
[0158] Step S2031, input the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the first reasoning step.
[0159] It can be understood that since the first large language model can output formatted reasoning steps, after inputting the first logical problem into the first large language model corresponding to any set of hyperparameters, use the specified step separator as the end symbol to sample different reasoning steps.
[0160] Step S2032, input the first logical problem and the multiple first reasoning steps corresponding to the first logical problem into the third large language model to obtain the target first reasoning step and the evaluation information of the target first reasoning step, where the logical reasoning ability of the third large language model is higher than that of the pre-trained large language model.
[0161] Among them, the evaluation information includes a score and guiding opinions.
[0162] Input the first logical problem into the first large language model corresponding to any set of hyperparameters in multiple sets of hyperparameters to obtain the first reasoning steps output by the first large language model corresponding to each set of hyperparameters. Then, for the first logical problem, multiple first reasoning steps are obtained.
[0163] Use the third large language model to screen multiple first reasoning steps to obtain the target first reasoning step. Exemplarily, 32 first reasoning steps are obtained, and the 32 first reasoning steps and the first logical problem are input into the third large language model to obtain eight target first reasoning steps.
[0164] The evaluation information of eight target first reasoning steps can also be obtained by using the third large language model.
[0165] Step S2033: Based on the target first reasoning step, determine the target hyperparameters corresponding to the target first reasoning step.
[0166] After determining the target first reasoning step, determine the hyperparameters of the first large language model that outputs the target first reasoning step as the target hyperparameters.
[0167] Step S2034: Input the first logical problem, the target first reasoning step, and the evaluation information of the target first reasoning step into the first large language model corresponding to the target hyperparameters to obtain the second reasoning step.
[0168] Among them, after obtaining the target first reasoning step, the long thinking chain can be extended based on the target first reasoning step to obtain step-level thinking chain data.
[0169] Step S2035: Use the third large language model to determine the evaluation information of the second reasoning step based on the target first reasoning step, the second reasoning step, and the first logical problem.
[0170] Among them, after obtaining the second reasoning step corresponding to each target first reasoning step, input the target first reasoning step, the second reasoning step, and the first logical problem into the third large language model to obtain the evaluation information of the second reasoning step. This evaluation information includes guiding opinions, which include guiding opinions on the second reasoning step and guiding opinions on the next reasoning step of the second reasoning step.
[0171] When using the third large language model to generate guiding opinions for the second reasoning step, the following prompt can be used:
[0172] "Please conduct a detailed analysis of the current step in the following reasoning chain:
[0173] 1. First, check whether this step is logically correct and reasonable, whether it conforms to the meaning of the question, and whether there are misunderstandings or deviations.
[0174] 2. Confirm whether this step clearly expresses the reasoning process and can effectively lead to the next step of reasoning.
[0175] 3. Put forward specific opinions on possible improvements, including how to optimize the expression, supplement missing information, or adjust the reasoning order, etc.
[0176] 4. Finally, based on the analysis of the current step, give specific guidance on the generation of the next step to ensure that the next step of reasoning can be closely connected and promote the solution of the problem.
[0177] Analyze the current steps in bullet points and give corresponding opinions.
[0178] Among them, the reasoning chain can include the reasoning steps obtained currently, arranged in the order of acquisition.
[0179] In step S2036, input the target first reasoning step, second reasoning step, first logical problem, and evaluation information of the second reasoning step into the first large language model corresponding to the target hyperparameter to obtain a third reasoning step.
[0180] In step S2037, iteratively execute using the third large language model to determine the evaluation information of the currently obtained reasoning step based on the first logical problem, the reasoning steps before the currently obtained reasoning step, and the currently obtained reasoning step. Input the first logical problem, the reasoning steps before the currently obtained reasoning step, the currently obtained reasoning step, and the evaluation information of the currently obtained reasoning step into the first language model corresponding to the target hyperparameter to obtain the next reasoning step until the logical reasoning ends and obtain the thinking chain data of the first logical problem.
[0181] It can be understood that according to the idea of the debate between two agents, continuously iterate and guide to obtain the final thinking chain data. Among them, the two agents are the first large language model and the third large language model.
[0182] It should be noted that the method for obtaining thinking chain data in this embodiment can also be applied to the acquisition of multi-modal long thinking chain data.
[0183] In step S204, based on the first logical problem, the thinking chain data of the first logical problem, and the evaluation information of the reasoning steps in the thinking chain data, adjust the model parameters of the first large language model to obtain a second large language model with preliminary logical reasoning ability and a reasoning step scoring model. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.
[0184] In step S205, obtain a second set of logical problems. Based on the second set of logical problems and the reasoning step scoring model, adjust the model parameters of the second large language model to obtain a target large language model with generalization ability, where the amount of logical problem data in the second set of logical problems is more than that in the first set of logical problems. For details, please refer to Figure 1 Step S105 of the embodiment shown, which will not be elaborated here.
[0185] It should be noted that steps S204 and S205 can be carried out alternately, which can not only improve the generalization ability but also have better logical reasoning ability in the corresponding technical field.
[0186] Further, it should be noted that the target large language model of this embodiment can be used orthogonally with the multi-path search technology to further improve the logical reasoning ability of the large language model, enabling the large language model to have the multi-path long-range thinking ability.
[0187] The large language model training method provided in this embodiment adjusts the model parameters of the pre-trained large language model to obtain a first large language model with formatted output, improving the interpretability and transparency of the model output, ensuring the format consistency and standardization of the model output results, and reducing misunderstandings or errors caused by inconsistent output formats.
[0188] The large language model training method provided in this embodiment inputs the first logical problem into the first large language model corresponding to any set of hyperparameters to obtain the first reasoning step, and then uses a third large language model with stronger logical reasoning ability to screen the first reasoning step to obtain the target first reasoning step and the evaluation information of the target first reasoning step. While ensuring the diversity of the generated thought chain data, the computational amount of generating the thought chain data is reduced.
[0189] In each iteration process, based on the currently obtained reasoning steps and their evaluation information, the subsequent reasoning steps are further adjusted and optimized to ensure the logical consistency, coherence, and accuracy of the finally output thought chain data, ensuring that the generated thought chain data has actions of making mistakes and correcting mistakes, thereby improving the logical reasoning ability of the second large language model and improving the processing effect of complex logical problems.
[0190] In some optional implementation manners, the above step S2012 includes:
[0191] Step a1: Input the training sample into the pre-trained large language model to obtain a first output result, where the first output result represents the probability value of the label marked by the pre-trained large language model corresponding to the training sample.
[0192] Step a2: Based on the first output result, determine a first loss value, where the first loss value includes the loss value of the format symbol.
[0193] Step a3: Based on the first loss value, adjust the model parameters of the pre-trained large language model until the adjustment termination condition is reached to obtain a first large language model with formatted output.
[0194] Among them, the gradient descent algorithm is used to update the model parameters of the pre-trained large language model to minimize the first loss value.
[0195] Based on the first loss value, adjusting the model parameters of the pre-trained large language model until the adjustment termination condition is reached to obtain a first large language model with formatted output includes:
[0196] Based on the first loss value, adjust the model parameters of the pre-trained large language model until the first loss value is not greater than the first loss value threshold, and obtain the first large language model with formatted output.
[0197] It can be understood that the adjustment termination condition can be that the first loss value is not greater than the first loss value threshold. Among them, the first loss value threshold is set by the technical personnel and is not specifically limited here.
[0198] The large language model training method provided in this embodiment can significantly improve the accuracy and robustness of the model by continuously optimizing the model parameters to gradually reduce the first loss value until it is not greater than the set first loss value threshold.
[0199] The large language model training method provided in this embodiment obtains the loss value of the format symbol, and based on the loss value of the format symbol, adjusts the model parameters of the pre-trained large language model to obtain the first large language model with formatted output, improving the interpretability and transparency of the model output, ensuring the format consistency and standardization of the model output results, and reducing misunderstandings or errors caused by inconsistent output formats.
[0200] In some alternative embodiments, step a2 includes:
[0201] Step a21, determine the position marked by the format symbol based on the reasoning steps corresponding to the logical problems in the training samples.
[0202] Marked as Token, which refers to the information of a single word understood by the computer and can be understood as a standard word or Chinese word. The large language model outputs are all individual tokens.
[0203] Exemplarily, the format symbol can be symbols such as "The first step:", "The second step:", " " and other symbols in terms of format. The format symbol is marked as the token corresponding to the format symbol.
[0204] Step a22, based on the position marked by the format symbol, determine the probability value that the token corresponding to the position marked by the format symbol is the token corresponding to the format symbol from the first output result.
[0205] Step a23, determine the first loss value based on the probability value that the token corresponding to the position marked by the format symbol is the token corresponding to the format symbol.
[0206] Specifically, determine the first loss value by inputting the probability value that the token corresponding to the position marked by the format symbol is the token corresponding to the format symbol into the first loss function.
[0207] Among them, the first loss function is:
[0208]
[0209] Among them, is the first loss value, and k is the position marked by the format symbol. is for the first (i - 1) tokens in the first output result marked as , and when the model parameters of the pre-trained large language model are 1, the probability value that the i-th token in the first output result is the token corresponding to the format symbol.
[0210] That is to say, is the probability value that the token corresponding to the position of the format symbol in the first output result is the token corresponding to the format symbol.
[0211] It should be noted that this step only fine-tunes the output format of the pre-trained large language model. Therefore, when calculating the loss value, only the loss value of the format symbol is calculated.
[0212] This first loss function can measure the difference between the token corresponding to the format symbol in the output result of the pre-trained large language model and the token corresponding to the format symbol in the actual reasoning step.
[0213] In the large language model training method provided in this embodiment, the first loss value is determined by the probability value that the token corresponding to the position of the format symbol marked in the first output result is the token corresponding to the format symbol. Based on the first loss value, the model parameters of the pre-trained large language model are adjusted to obtain the first large language model with formatted output, ensuring the formatting and consistency of the model output.
[0214] In some alternative embodiments, the above step S2032 includes:
[0215] Step b1, using the third large language model, based on the first logical problem, screening the multiple first reasoning steps corresponding to the first logical problem from the dimensions of correctness, uncertainty, and repeatability to obtain the target first reasoning steps.
[0216] Among them, since the first reasoning step directly affects the analysis of the logical problem and the subsequent thinking chain, in order to ensure the diversity of the thinking chain, the importance difference sampling technique is adopted. Using the third large language model, based on the first logical problem, screening the multiple first reasoning steps corresponding to the first logical problem from the dimensions of correctness, uncertainty, and repeatability to obtain the target first reasoning steps.
[0217] Specifically, the third large language model screens out the target first reasoning steps through the method of prompt engineering. Exemplarily, the prompt can be:
[0218] "Please select 8 pieces of data that meet the criteria from the following 32 pieces of data. When making the selection, please judge according to the following 3 indicators:
[0219] 1. Correctness: Select data with no errors, logical consistency, and clear information;
[0220] 2. Uncertainty: Prioritize data with uniqueness, inspiration, or additional room for interpretation to avoid selecting content with too high certainty;
[0221] 3. Repetition: Avoid duplicate or similar content and prioritize data with independent and non-repetitive properties.
[0222] The 8 pieces of data selected should be representative in terms of overall content and quality, while avoiding data that does not meet the requirements. Please strictly complete the selection according to these three criteria and output a list of 8 pieces of data."
[0223] It can be understood that the data in the prompt is the first reasoning step.
[0224] The large language model training method provided in this embodiment screens multiple first reasoning steps corresponding to the first logical problem from the dimensions of correctness, uncertainty, and repetition by using a third large language model based on the first logical problem, obtains the target first reasoning step, and removes the first reasoning steps that do not meet the requirements, ensuring the accuracy of the obtained thought chain data.
[0225] In some alternative embodiments, the above step S204 includes:
[0226] Step c1, input the first logical problem and the thought chain data of the first logical problem into the first large language model to obtain a second output result, where the second output result represents the probability value of the labels output by the first large language model marked as the labels corresponding to the first logical problem and the thought chain data of the first logical problem.
[0227] Among them, one first logical problem and the corresponding thought chain data of the first logical problem are a training sample.
[0228] Step c2, based on the second output result, determine a second loss value, where the second loss value includes the loss value of the reasoning steps of the thought chain data.
[0229] Among them, input the second output result into the second loss function to obtain the second loss value.
[0230] It should be noted that the second loss function is similar to the first loss function, but calculates the loss value of the reasoning steps of the thought chain data.
[0231] Step c3: Based on the second loss value, adjust the model parameters of the first large language model to obtain a second large language model with preliminary logical reasoning ability.
[0232] Step c4: Input the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data into the first large language model to obtain a third output result. The third output result represents the probability values of the tokens output by the first large language model that are marked as corresponding to the first logical problem, the thought chain data of the first logical problem, and the evaluation information of the reasoning steps in the thought chain data.
[0233] Among them, the evaluation information of the reasoning steps in the thought chain data can be placed after the corresponding reasoning steps of the thought chain data of the first logical problem.
[0234] The first logical problem, a thought chain data corresponding to the first logical problem, and the evaluation information of the reasoning steps in the thought chain data form a training sample.
[0235] Step c5: Based on the third output result, determine the third loss value. The third loss value includes the loss value of the scoring of the reasoning steps.
[0236] Among them, the evaluation information includes scoring. In the case where the scoring of the reasoning steps is placed after each reasoning step, calculate the loss value of the scoring after each reasoning step, which is the third loss value.
[0237] Among them, input the third output result into the third loss function to determine the third loss value. It should be noted that the third loss function is similar to the first loss function, but calculates the loss value of the scoring of the reasoning steps.
[0238] Step c6: Based on the third loss value, adjust the model parameters of the first large language model to obtain a reasoning step scoring model. Among them, the evaluation information includes scoring.
[0239] The large language model training method provided in this embodiment improves the logical reasoning ability of the model and the accuracy of the reasoning steps by calculating the loss value of the reasoning steps of the thought chain data and the loss value of the scoring of the reasoning steps, and adjusting the model parameters of the first large language model based on the loss value of the reasoning steps of the thought chain data and the loss value of the scoring of the reasoning steps respectively.
[0240] In some alternative embodiments, the above step S205 includes:
[0241] Step d1: Create multiple environments that share the second large language model.
[0242] Among them, the second largest language model has preliminary logical reasoning ability, but its generalization ability is poor. To improve the generalization ability of the model, the Proximal Policy Optimization (PPO) method with multi-environment asynchrony is used to enhance the generalization ability of the second largest language model and obtain the target large language model. The PPO method is a reinforcement learning method.
[0243] Among them, the process of using the PPO method to enhance the generalization ability of the second largest language model includes:
[0244] Step 1: Create multiple environments (Envs) with multiple threads. There is a shared agent in each environment (Env). Based on any second logical problem in the second logical problem set, the agent in each environment infers the first inference step, recorded as action A1, that is, action(A1), and updates the state (observation) of the environment to the first logical problem and the first inference step. In this embodiment, the agent is the second largest language model.
[0245] Step 2: Create a runtime instance for reinforcement learning training. The instance mainly includes: the multiple created environments, the agent, the buffer space, the Parametric Reward Model (PRM), and the PPO reinforcement learning trainer. Among them, the PRM is the aforementioned inference step scoring model.
[0246] Step 3: Reset the environment, given a second logical problem, and start a reinforcement learning process.
[0247] Step 4: Start an inference of the second largest language model to obtain action(A1). This action(A1) is the decode tokens or tokens, and the probability value (logits) of action(A1). Among them, the probability value of action(A1) is obtained after inputting the state of the environment and action(A1) into the LLM.
[0248] Step 5: Feed the action inferred by the second largest language model and the state of the environment into the PRM to obtain the reward value, that is, the score of the action.
[0249] Step 6: Return to execute Step 1 to obtain the state of the new environment.
[0250] Step 7: Insert the reward value, probability value, and action into the buffer for later use.
[0251] Repeat steps 4 to 7 until all simulations in all environments are completed. The sign that all simulations in all environments are completed is that the environment reaches the completed (done) state or the simulation depth (the threshold of the number of inference steps) reaches the preset depth threshold. Exemplarily, the preset depth threshold can be 8.
[0252] Step 8, Training phase, use the data stored in the buffer, i.e., the value of each inference step, etc., and perform learning and training using the PPO strategy. The data in the buffer includes: (1) obs_batch (state data, i.e., logical problems); (2) action_batch (actions in the state, i.e., inference steps); (3) log_prob_batch (confidence value of the action, i.e., probability value of the inference step); (4) reward_batch (reward value of the current state, i.e., score of the inference step); (5) action_tokens_batch (identifier of the action, position of the inference step in the preset dictionary).
[0253] Step 9, During the training process, update the model parameters of the agent, i.e., update the model parameters of the second large language model to obtain the target large language model.
[0254] Step d2, In each environment, for any second logical problem in the second logical problem set, input the second logical problem into the second large language model to obtain the first inference step of the second logical problem.
[0255] Corresponding to the aforementioned inference of starting the second large language model once, obtain action (A1).
[0256] Step d3, Obtain the probability value of the first inference step.
[0257] Step d4, Input the first inference step and the second logical problem into the inference step scoring model to obtain the score of the first inference step.
[0258] Step d5, Input the second logical problem and the first inference step into the second large language model to obtain the second inference step of the second logical problem.
[0259] Step d6, Obtain the probability value of the second inference step.
[0260] Step d7, Input the second inference step, the second logical problem, and the first inference step into the inference step scoring model to obtain the score of the second inference step.
[0261] Step d8, iteratively execute inputting the second logical problem and the inference steps before the inference step to be obtained into the second large language model, obtain the inference step to be obtained, acquire the probability value of the inference step to be obtained, input the second logical problem, the inference step to be obtained, and the inference steps before the inference step to be obtained into the inference step scoring model, obtain the score of the inference step to be obtained, until the logical reasoning of the environment for the second logical problem ends.
[0262] Step d9, save the second logical problem, the inference steps of the second logical problem, the probability value of the inference step, and the score of the inference step during the logical reasoning process.
[0263] Step d10, based on the saved second logical problem, the inference steps of the second logical problem, the probability value of the inference step, and the score of the inference step, adjust the model parameters of the second large language model to obtain a target large language model with generalization ability.
[0264] Among them, the relevant descriptions of steps d1 to d10 can refer to steps 1 to 9 in the process of using the PPO method to enhance the generalization ability of the second large language model.
[0265] The large language model training method provided in this embodiment, by creating multiple environments, enables the model to be fully trained in different environments, enhances the generalization ability of the model, enables the model to provide reliable and accurate answers when facing unseen logical problems, and improves its applicability in various scenarios.
[0266] In some alternative embodiments, the above step d3 includes:
[0267] Step d31, input the second logical problem and the first inference step into the second large language model to obtain the probability value of the first inference step.
[0268] It can be understood that the calculation method of the probability values of the remaining inference steps is similar to that of the probability value of the first inference step, and will not be elaborated here.
[0269] The large language model training method provided in this embodiment, by obtaining the probability value of the first inference step and guiding the model to be optimized based on the probability value of the first inference step, enhances the logical reasoning ability of the model.
[0270] In some alternative embodiments, the above step d10 includes:
[0271] Step d101: Based on the saved second logical problem, the reasoning steps of the second logical problem, and the probability values of the reasoning steps, determine the adjacent reasoning steps of the second logical problem under the reference model and the probability value of each reasoning step in the adjacent reasoning steps, where the reference model is the second largest language model without model parameter adjustment.
[0272] Exemplarily, input the second logical problem into the reference model to obtain the reasoning steps of the second logical problem under the reference model. The first reasoning step and the second reasoning step in the reasoning steps of the second logical problem under the reference model are adjacent reasoning steps, and the second reasoning step and the third reasoning step are adjacent reasoning steps. It should be noted that the first and second in the first reasoning step and the second reasoning step represent the order of obtaining the reasoning steps.
[0273] It can be understood that the saved reasoning steps of the second logical problem and the probability values of the reasoning steps are obtained by inputting the second logical problem into the reference model. That is to say, the saved reasoning steps of the second logical problem and the probability values of the reasoning steps are the reasoning steps of the second logical problem under the reference model and the probability values of the reasoning steps.
[0274] Step d102: Based on the scores of the saved reasoning steps, determine the score of each reasoning step in the adjacent reasoning steps.
[0275] Among them, the scores of the saved reasoning steps are obtained through the reasoning step scoring model.
[0276] Step d103: Based on the adjacent reasoning steps of the second logical problem under the reference model, the probability value of each reasoning step in the adjacent reasoning steps, and the score of each reasoning step in the adjacent reasoning steps, determine the fourth loss value.
[0277] Step d104: Based on the fourth loss value, adjust the model parameters of the second largest language model to obtain a target large language model with generalization ability.
[0278] The large language model training method provided in this embodiment determines the fourth loss value based on the adjacent reasoning steps of the second logical problem under the reference model, the probability value of each reasoning step in the adjacent reasoning steps, and the score of each reasoning step in the adjacent reasoning steps. Based on the fourth loss value, the model parameters of the second largest language model are adjusted to obtain a target large language model with generalization ability, enhancing the logical reasoning ability and generalization ability of the model.
[0279] In some optional implementation manners, the above step d103 includes:
[0280] Step d1031: Obtain the probability value of each reasoning step in the adjacent reasoning steps of the second logical problem under the second largest language model corresponding to the current model parameters.
[0281] Among them, by inputting the second logical problem and the corresponding reasoning steps into the second large language model, the probability value of each reasoning step is obtained, and then the probability value of each reasoning step in adjacent reasoning steps is determined.
[0282] Step d1032, for any one of the reasoning steps in adjacent reasoning steps, based on the probability value of this reasoning step in the adjacent reasoning steps of the second large language model corresponding to the current model parameters for the second logical problem and the probability value of this reasoning step in the adjacent reasoning steps of the second logical problem in the reference model, determine the divergence between the reference model and the second large language model corresponding to the current model parameters under the reasoning step, so as to obtain the divergence between the reference model and the second large language model corresponding to the current model parameters under each reasoning step in adjacent reasoning steps.
[0283] Step d1033, based on the score of each reasoning step in adjacent reasoning steps and the divergence between the reference model and the second large language model corresponding to the current model parameters under each reasoning step in adjacent reasoning steps, determine the fourth loss value.
[0284] Specifically, input the score of each reasoning step in adjacent reasoning steps and the divergence between the reference model and the second large language model corresponding to the current model parameters under each reasoning step in adjacent reasoning steps into the fourth loss function to obtain the fourth loss value.
[0285] Among them, the fourth loss function is:
[0286]
[0287] Among them, is the fourth loss value, is the second large language model with the current model parameters being ; is to find the maximum value, is the divergence, E is the expectation, is the second logical problem, is the score of the i-th reasoning step in adjacent reasoning steps, is the reasoning step output by the reference model, is the hyperparameter, is the probability value of the i-th reasoning step in the adjacent reasoning steps of the second logical problem in the reference model, is the probability value of the i-th reasoning step in the adjacent reasoning steps of the second logical problem in the second large language model with the current model parameters being ; is the reference model, is the model input corresponding to the i-th reasoning step in the adjacent reasoning steps of the second logical problem 1 is the first inference step in adjacent inference steps, 2 is the second inference step in adjacent inference steps, and the reference model is the second largest language model without model parameter adjustment.
[0288] It can be understood that when the adjacent inference steps are the first inference step and the second inference step, when i = 1, B is the second logical problem, and when i = 2, B is the second logical problem and the first inference step.
[0289] That is to say, is the divergence of the second largest language model corresponding to the reference model and the current model parameters at the i-th inference step in adjacent inference steps. Among them, the value of B changes according to the value of i, that is changes according to the value of i.
[0290] The large language model training method provided in this embodiment determines the fourth loss value through the score of each inference step in adjacent inference steps and the divergence of the second largest language model corresponding to the reference model and the current model parameters at each inference step in adjacent inference steps. Based on the fourth loss value, the model parameters of the second largest language model are adjusted to obtain a target large language model with generalization ability, which helps to guide the model to learn a more reasonable and accurate inference process, thereby improving the performance of the model in logical reasoning tasks and making the output inference steps more logical and in line with the actual situation.
[0291] In some alternative embodiments, the above step d1033 includes:
[0292] Step e1, determining a first reward value based on the score of the first inference step in adjacent inference steps and the divergence of the second largest language model corresponding to the reference model and the current model parameters at the first inference step.
[0293] Among them, for the fourth loss function, when i = 1, by calculating - , the first reward value is determined.
[0294] Step e2, determining the second reward value of the second inference step in adjacent inference steps based on the scores of each inference step in adjacent inference steps.
[0295] Specifically, by inputting the scores of each inference step in adjacent inference steps into the reward loss function, the second reward value is obtained.
[0296] Among them, the reward loss function is:
[0297]
[0298] Among them, is the second reward value, is a hyperparameter, is the score of the second inference step in adjacent inference steps, is the score of the first inference step in adjacent inference steps.
[0299] Step e3, based on the score of the second inference step in adjacent inference steps, the divergence between the reference model and the second - largest language model corresponding to the current model parameters at the second inference step, and the second reward value of the second inference step in adjacent inference steps, determine the third reward value.
[0300] Among them, for the fourth loss function, when i = 2, update the in the fourth loss function to { + } to calculate the third reward value.
[0301] That is, by calculating - , determine the third reward value.
[0302] Step e4, based on the first reward value and the third reward value, determine the fourth loss value.
[0303] Among them, adjacent inference steps include the first inference step and the second inference step. The first inference step is the inference step with a higher acquisition order in adjacent inference steps, and the second inference step is the inference step with a lower acquisition order in adjacent inference steps.
[0304] Specifically, by inputting the first reward value and the third reward value into the fourth loss function, determine the fourth loss value.
[0305] This fourth loss function is an improved PPO policy loss function, which can reward the reflection behavior and correction behavior between inference steps.
[0306] The large - language model training method provided in this embodiment rewards the self - improvement behavior that occurs during the model's inference process, encourages the model to continuously optimize in subsequent inferences, attempts to generate better inference steps, and thus gradually improves the overall performance and inference quality of the model.
[0307] In this embodiment, a large - language model training method is provided, which can be used in mobile terminals, such as servers, central processing units, etc. Figure 3 is the flowchart of the large - language model training method according to the embodiment of the present invention, as Figure 3 shown. This process includes the following steps:
[0308] Step S301, format and fine - tune the large - language model.
[0309] Among them, the fine-tuning of the LLM format means obtaining the first large language model with output formatting. For details, please refer to the description of the foregoing step S201, which will not be elaborated here.
[0310] Step S302: Data collection for the first inference step of the large language model.
[0311] Among them, the data collection of the LLM step1 (the first inference step) means inputting the first logical problem and multiple first inference steps corresponding to the first logical problem into the third large language model to obtain the target first inference step. For details, please refer to the description of the foregoing step S2032, which will not be elaborated here.
[0312] Step S303: Data collection of the long thinking chain for the second inference step and subsequent inference steps based on multiple agents.
[0313] Among them, for the data collection of the long thinking chain of the step2-stepN (the second inference step and subsequent inference steps) based on multiple agents, please refer to the descriptions of the foregoing steps S2034 to S2037, which will not be elaborated here. The multiple agents include the first large language model and the third large language model.
[0314] Step S304: Fine-tuning of the large language model content (long thinking chain).
[0315] Among them, the fine-tuning of the LLM content (long thinking chain) means adjusting the model parameters of the first large language model based on the first logical problem, the thinking chain data of the first logical problem, and the evaluation information of the inference steps in the thinking chain data, so as to obtain the second large language model with preliminary logical reasoning ability and the inference step scoring model. For details, please refer to the description of the foregoing step S204, which will not be elaborated here.
[0316] Step S305: Generalization training of reinforcement learning.
[0317] Among them, the generalization training of reinforcement learning means adjusting the model parameters of the second large language model based on the second logical problem set and the inference step scoring model to obtain the target large language model with generalization ability. For details, please refer to the description of the foregoing step S205, which will not be elaborated here.
[0318] It is understandable that when cultivating the logical reasoning ability of an LLM, it is crucial to obtain data containing incorrect steps and corresponding modifications, as well as correct long thinking chain data. The large language model training method provided in this embodiment obtains thinking chain data by using the importance difference sampling method and leveraging multi-agent to expand the step-level long thinking chain data, and fine-tunes the first large language model based on the thinking chain data to obtain a second large language model with preliminary logical reasoning ability. The preliminary logical reasoning ability is the long-range thinking ability that is logically consistent, thinks step by step, and self-reflects and modifies. The second large language model has long-range thinking ability within a certain knowledge domain, but due to data limitations, its generalization ability and step reflection ability are poor. By using the PPO method, adding step-level error correction to the loss function constraint, and enhancing and generalizing the long-range thinking ability in the reinforcement learning environment, a target large language model is obtained. The target large language model has strong logical reasoning ability and improves the processing effect of complex logical problems.
[0319] In this embodiment, a method for processing logical problems is provided, which can be used in mobile terminals such as servers and central processing units. The method for processing logical problems includes:
[0320] Obtain the logical problem to be processed.
[0321] Input the logical problem to be processed into the target large language model trained by using the large language model training method shown in any of the above embodiments to obtain the thinking chain data corresponding to the logical problem to be processed.
[0322] The method for processing logical problems provided in this embodiment realizes the efficient and accurate processing of logical problems by using the target large language model trained by using the large language model training method shown in any of the above embodiments to process the logical problem to be processed.
[0323] The embodiment of the present invention also provides a computer device. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 4As shown, the computer device includes: one or more processors 401, a memory 402, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 4 In [the figure], a single processor 401 is taken as an example.
[0324] The processor 401 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 401 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0325] Among them, the memory 402 stores instructions executable by at least one processor 401, so that at least one processor 401 executes the method shown in the above embodiments.
[0326] The memory 402 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 402 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 402 can optionally include a memory remotely set relative to the processor 401, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0327] The memory 402 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 402 can also include a combination of the above types of memories.
[0328] The computer device further includes a communication interface 403 for the computer device to communicate with other devices or communication networks.
[0329] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0330] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0331] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A large language model training method, characterized in that: The method comprises: Get the first language model for output formatting; Obtain multiple sets of hyperparameters for a first logic problem set and a first language model; For any first logical problem in the first logical problem set, the first logical problem is input into a first large language model corresponding to any set of hyperparameters to obtain thinking chain data of the first logical problem, the thinking chain data including a plurality of reasoning steps, and the large language model generates a current reasoning step based on the first logical problem, a previous reasoning step and evaluation information of the previous reasoning step, wherein the evaluation information includes a score and guidance; Based on the first logical question, the thinking chain data of the first logical question and the evaluation information of the reasoning steps in the thinking chain data, the model parameters of the first language model are adjusted to obtain the second language model and the reasoning step scoring model with preliminary logical reasoning ability; Obtaining a second logical question set, and adjusting model parameters of a second large language model based on the second logical question set and the reasoning step scoring model to obtain a target large language model with generalization capability, wherein the amount of logical question data of the second logical question set is greater than the amount of logical question data of the first logical question set; Obtaining a reasoning step scoring model, including: inputting a first logical question, a thinking chain data of the first logical question, and evaluation information of the reasoning steps in the thinking chain data into a first large language model, obtaining a third output result, wherein the third output result represents a probability value of a mark output by the first large language model corresponding to the first logical question, the thinking chain data of the first logical question, and the evaluation information of the reasoning steps in the thinking chain data; Based on the third output result, determine a third loss value, the third loss value comprising a loss value of the score of the reasoning step; Based on the third loss value, the model parameters of the first language model are adjusted to obtain an inference step scoring model.
2. The method according to claim 1, characterized in that The first language model for obtaining output formatting includes: Acquire a data set, wherein training samples in the data set include logic problems and reasoning steps corresponding to the logic problems; Based on the data set, model parameters of the pre-trained large language model are adjusted to obtain a first large language model with formatted output.
3. The method according to claim 2, characterized in that The step of adjusting the model parameters of the pre-trained large language model based on the data set to obtain the first large language model with output formatting includes: Inputting the training sample into the pre-trained large language model to obtain a first output result, wherein the first output result represents a probability value of a label output by the pre-trained large language model being a label corresponding to the training sample; Based on the first output result, determining a first loss value, the first loss value comprising a loss value of a format symbol; Based on the first loss value, the model parameters of the pre-trained large language model are adjusted until an adjustment termination condition is reached, thereby obtaining a first large language model with output formatting.
4. The method according to claim 3, characterized in that: The determining a first loss value based on the first output result includes: Determining the position of the format symbol mark based on the reasoning steps corresponding to the logic problem in the training sample; Based on the position of the format symbol mark, determining from the first output result a probability value that the mark corresponding to the position of the format symbol mark is the mark corresponding to the format symbol; A first loss value is determined based on a probability value that the mark corresponding to the position of the format symbol mark is the mark corresponding to the format symbol.
5. The method according to claim 2, characterized in that: The step of inputting the first logical question into a first language model corresponding to any set of hyperparameters to obtain the thought chain data of the first logical question includes: Inputting the first logical question into a first language model corresponding to any set of hyperparameters to obtain a first reasoning step; Inputting the first logic question and a plurality of first reasoning steps corresponding to the first logic question into a third large language model to obtain a target first reasoning step and evaluation information of the target first reasoning step, wherein the logic reasoning ability of the third large language model is higher than the logic reasoning ability of the pre-trained large language model; Based on the target first reasoning step, determining a target hyperparameter corresponding to the target first reasoning step; Inputting the first logical question, the target first reasoning step, and evaluation information of the target first reasoning step into a first large language model corresponding to the target hyperparameter to obtain a second reasoning step; Determining evaluation information of the second reasoning step based on the target first reasoning step, the second reasoning step and the first logical question using the third language model; Inputting the target first reasoning step, the second reasoning step, the first logical question, and the evaluation information of the second reasoning step into the first large language model corresponding to the target hyperparameter to obtain a third reasoning step; Iterative execution utilizes the third largest language model to determine evaluation information of the currently obtained reasoning step based on the first logical problem, the reasoning step before the currently obtained reasoning step, and the currently obtained reasoning step, and inputs the first logical problem, the reasoning step before the currently obtained reasoning step, the currently obtained reasoning step, and the evaluation information of the currently obtained reasoning step into the first language model corresponding to the target hyperparameter to obtain the next reasoning step, until the logical reasoning is completed, and the thinking chain data of the first logical problem is obtained.
6. The method according to claim 5, characterized in that The step of inputting the first logical question and a plurality of first reasoning steps corresponding to the first logical question into a third language model to obtain a target first reasoning step includes: By using the third language model and based on the first logical problem, multiple first reasoning steps corresponding to the first logical problem are screened from the dimensions of correctness, uncertainty and repeatability to obtain the target first reasoning step.
7. The method according to claim 1, characterized in that Based on the first logical question, the thinking chain data of the first logical question and the evaluation information of the reasoning steps in the thinking chain data, the model parameters of the first language model are adjusted to obtain a second language model with preliminary logical reasoning ability, including: Inputting the first logical question and the thinking chain data of the first logical question into the first large language model to obtain a second output result, wherein the second output result represents a probability value of a mark output by the first large language model that is marked as corresponding to the first logical question and the thinking chain data of the first logical question; Based on the second output result, determining a second loss value, wherein the second loss value includes a loss value of a reasoning step of the thought chain data; Based on the second loss value, the model parameters of the first language model are adjusted to obtain a second language model with preliminary logical reasoning capabilities.
8. The method according to claim 1, characterized in that The adjusting the model parameters of the second large language model based on the second logical question set and the reasoning step scoring model to obtain a target large language model with generalization capability includes: Creating multiple environments, where the multiple environments share the second largest language model; In each environment, for any second logical problem in the second logical problem set, input the second logical problem into the second largest language model to obtain a first reasoning step of the second logical problem; Obtaining a probability value of the first reasoning step; Inputting the first reasoning step and the second logic question into the reasoning step scoring model to obtain a score for the first reasoning step; Inputting the second logical problem and the first reasoning step into the second largest language model to obtain a second reasoning step for the second logical problem; Obtaining a probability value of the second reasoning step; Inputting the second reasoning step, the second logical question and the first reasoning step into the reasoning step scoring model to obtain a score for the second reasoning step; Iteratively execute inputting the second logical question and the reasoning step before the reasoning step to be obtained into the second largest language model, obtaining the reasoning step to be obtained, acquiring the probability value of the reasoning step to be obtained, inputting the second logical question, the reasoning step to be obtained, and the reasoning step before the reasoning step to be obtained into the reasoning step scoring model, obtaining the score of the reasoning step to be obtained, until the logical reasoning of the environment for the second logical question is completed; The second logical question, the reasoning steps of the second logical question, the probability values of the reasoning steps, and the scores of the reasoning steps in the logical reasoning process are saved; Based on the saved second logical problem, the reasoning steps of the second logical problem, the probability values of the reasoning steps and the scores of the reasoning steps, the model parameters of the second large language model are adjusted to obtain a target large language model with generalization ability.
9. The method according to claim 8, characterized in that The obtaining the probability value of the first reasoning step includes: The second logical question and the first reasoning step are input into the second large language model to obtain a probability value of the first reasoning step.
10. The method according to claim 8, characterized in that The method of adjusting the model parameters of the second large language model based on the saved second logical problem, the reasoning steps of the second logical problem, the probability values of the reasoning steps, and the scores of the reasoning steps to obtain a target large language model with generalization ability includes: Determine adjacent reasoning steps of the second logical problem under a reference model and the probability value of each of the adjacent reasoning steps based on the saved second logical problem, the reasoning steps of the second logical problem, and the probability value of the reasoning steps, wherein the reference model is the second largest language model without model parameter adjustment; Determining a score for each of the adjacent reasoning steps based on the saved scores of the reasoning steps; Determine a fourth loss value based on adjacent reasoning steps of the second logical problem under the reference model, a probability value of each reasoning step in the adjacent reasoning steps, and a score of each reasoning step in the adjacent reasoning steps; Based on the fourth loss value, the model parameters of the second large language model are adjusted to obtain a target large language model with generalization capability.
11. The method according to claim 10, characterized in that The determining of the fourth loss value based on the adjacent reasoning steps of the second logical problem under the reference model, the probability value of each reasoning step in the adjacent reasoning steps, and the score of each reasoning step in the adjacent reasoning steps comprises: Obtaining a probability value of each reasoning step in adjacent reasoning steps of the second logical problem under the second largest language model corresponding to the current model parameters; For any reasoning step in adjacent reasoning steps, based on the probability value of the reasoning step in the adjacent reasoning step of the second logical problem under the second largest language model corresponding to the current model parameters and the probability value of the reasoning step in the adjacent reasoning step of the second logical problem under the reference model, determine the divergence of the reference model and the second largest language model corresponding to the current model parameters under the reasoning step, so as to obtain the divergence of the reference model and the second largest language model corresponding to the current model parameters under each reasoning step in the adjacent reasoning steps; A fourth loss value is determined based on the score of each reasoning step in the adjacent reasoning steps and the divergence of the reference model and the second largest language model corresponding to the current model parameters at each reasoning step in the adjacent reasoning steps.
12. The method according to claim 11, characterized in that The determining of a fourth loss value based on a score of each reasoning step in the adjacent reasoning steps and a divergence between the reference model and the second largest language model corresponding to the current model parameters at each reasoning step in the adjacent reasoning steps comprises: Determining a first reward value based on a score of a first reasoning step in the adjacent reasoning steps and a divergence between the reference model and a second largest language model corresponding to the current model parameters under the first reasoning step; determining a second reward value for a second reasoning step in the adjacent reasoning steps based on the score of each reasoning step in the adjacent reasoning steps; determining a third reward value based on a score of a second reasoning step in the adjacent reasoning steps, a divergence of a second largest language model corresponding to the reference model and the current model parameters under the second reasoning step, and a second reward value of the second reasoning step in the adjacent reasoning steps; determining the fourth loss value based on the first reward value and the third reward value; The adjacent reasoning steps include a first reasoning step and a second reasoning step.
13. A method for processing a logic problem, characterized in that: include: Get pending logic issues; The logic problem to be processed is input into a target large language model trained by the large language model training method according to any one of claims 1 to 12 to obtain the thought chain data corresponding to the logic problem to be processed.
14. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the large language model training method described in any one of claims 1 to 12 or the logic problem processing method described in claim 13 by executing the computer instructions.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the large language model training method described in any one of claims 1 to 12 or the logic problem processing method described in claim 13.
16. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the large language model training method described in any one of claims 1 to 12 or the logic problem processing method described in claim 13.
Citation Information
Patent Citations
Large language model training method and device
CN118036757A
Large model training method, device and equipment and question answering method, device and equipment based on large model
CN118469019A
Big language model logical reasoning method and device, equipment, storage medium and product
CN119378691A