Reward model-based reinforcement learning for performing inference tasks

By combining language models and reward models, the problem of difficulty in performing complex inference tasks in the prior art is solved, and accurate and interpretable inference trajectories are generated in real-world environments, improving the practicality and security of the task.

CN119998819APending Publication Date: 2025-05-13GDM HOLDING LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380069293.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-28
Filing Date
2023-09-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize machine learning models, especially neural networks, to perform complex inference tasks, especially in real-world environments, and lacks effective methods to generate human-interpretable inference trajectories.

Method used

Using a combination of language model and reward model, the language model is trained to generate responses to input queries, and the reward model is used to guide the reinforcement learning of the language model to improve the accuracy and interpretability of the inference steps.

Benefits of technology

A language model for performing inference tasks in a real-world environment is realized, an accurate and interpretable inference trajectory is generated, and the practicality and security of mechanical agents and computer systems in real-world tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998819A_ABST
    Figure CN119998819A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a language model for performing inference tasks. The system obtains a plurality of training examples. Each training example includes a respective sample query text sequence characterizing a respective sample query and a respective reference response text sequence including a reference final answer to the respective sample query. The system trains a reward model over a plurality of training examples. The reward model is configured to receive an input comprising a sequence of query text characterizing a query and one or more inference steps that have been generated in response to the query, and process the input to compute a reward score indicating a degree of success of the one or more inference steps in generating a correct final answer to the query. The system trains a language model using the trained reward model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 377,532, filed on September 28, 2022, the disclosure of which is hereby incorporated by reference in its entirety. Technical Field

[0003] This specification relates to performing inference tasks using machine learning models such as neural networks. Background Art

[0004] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict outputs for received inputs. Some neural networks include one or more hidden layers in addition to the output layer. The output of each hidden layer is used as the input for the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates an output from the received input based on the current values ​​of the corresponding set of parameters.

[0005] The parameters of the machine learning model can be determined by a training process based on training data including one or more training examples. For example, a neural network can be trained by updating network parameters including, for example, weights and bias coefficients of a network layer of the neural network. Summary of the invention

[0006] This specification describes methods, computer systems, and apparatus, including computer programs encoded on computer storage media, for performing reasoning tasks using language models. Specifically, the reasoning task includes providing a response to an input query. More specifically, the response may include (1) a final answer to the query and (2) a reasoning trace that includes a sequence of one or more reasoning steps to produce the final answer to the query.

[0007] Language models and language generation neural networks can perform a variety of tasks, such as translation tasks (assuming that the training corpus includes words in different languages) and arithmetic tasks, as well as many other tasks. Related to reasoning tasks, language models can be used to generate responses to queries (e.g., questions). Responses to queries can be provided as text sequences, for example, in the form of yes / no binary answers, answers that provide one or more numerical values, answers that identify one or more objects or events, answers that identify planned routes or actions, etc. Responses to queries can be used in a variety of ways by humans or computer systems. For example, answers and reasoning trajectories may be useful in themselves, or may be used to provide warnings and / or control the movement of one or more robotic agents or objects, such as an autonomous vehicle.

[0008] In some embodiments, the input query can be a natural language query related to an environment, particularly a real-world environment. The output response can be a natural language reply or a natural language output statement that is also related to the environment. For example, the output response can provide information related to the environment, and in some embodiments, provide information related to an action to be taken in the environment or specify an action to be taken in the environment.

[0009] In one example, the system can be used to diagnose faults in a mechanical system operating in a real-world environment. Input queries can include, for example, observations about the mechanical system from one or more sensors (e.g., cameras, microphones, accelerometers, temperature sensors, etc.), and questions about the mechanical system. In some embodiments, observations can include sensed electronic signals, such as motor currents or temperature signals; and / or image or video data, such as images or video data from a camera or laser radar (LIDAR) sensor, such as data from a sensor of a mechanical agent or data from a sensor located separately from the mechanical agent in the environment. Observations and questions can be posed to generate a language representation of the query. For example, queries may include general questions such as “Given the measurements from sensors A, B, and C, is the system working correctly?” or “Given the measurements of A, B, and C, what is wrong with the system?” or specific requests such as “Is there a fault with component X?” The output response may provide the final answer to the question and the reasoning trail that led to the final answer.

[0010] In another example, the environment can be an educational environment, for example, the system can be deployed as part of an educational software program that assists a user in learning or practicing one or more corresponding skills. For example, the input query can be a math word problem, and the output response includes the correct solution to the word problem and the intermediate steps to reach the correct solution. The output response can be provided to a student as a demonstration of how to solve a math problem, or used as a grading tool to provide feedback on a student's practice.

[0011] In another example, the system can be used for natural language control of tasks in a real-world environment. That is, the input query can be related to the task, for example, it can include a request to perform the task. The output response can be used to control, for example, a mechanical system (which can be referred to as a mechanical agent) or a computer system for performing the task. As an example, the input query can include, for example, a high-level question from a human to perform a task, such as "What is the most cost-effective way to fabricate this component with our current equipment? (What is the most cost-effective way to manufacture this component using our current equipment?)" (in this case, the real-world environment may be a manufacturing facility). The output response includes the final answer to the specified plan, and the reasoning steps to determine the plan. The output response can be used to control one or more mechanical agents, such as a robotic arm, to perform the plan. For example, the output response can include one or more control signals for controlling one or more mechanical agents, such as the position, velocity, or force / torque / acceleration data of one or more joints of a robot or a part of another mechanical agent. In some embodiments, the input query can be related to a real-world task and can specify an observation of the real-world environment, and the output response can specify an action or route to be taken by a user when performing a task in a real-world environment. The system can provide output responses to a user, such as via a display or audio output device, to guide the user in performing a task in a real-world environment. Output responses generated using the methods described in this specification can greatly improve the practicality and safety of real-world tasks that can be performed by a mechanical agent, a computer system, or a human user in a real-world environment, particularly for tasks where generating an inaccurate response could have significant negative consequences. In particular, the output responses can be used to perform tasks in a principled and logical manner.

[0012] In some embodiments, the reasoning trace in the output response provides a human-interpretable explanation, which a user can use, for example, to 1) determine whether to perform the action specified in the final answer in the output response, or 2) access the explanation later when certain criteria are met to evaluate the performance of the language model or diagnose the cause of an error that occurred due to the execution of the action specified in the response generated by the language model. The human-interpretable explanation can, for example, include a sequence of logical steps presented in the form of natural language sentences that lead from the query to the response in a causal chain. Such human-interpretable explanations can overcome or mitigate some of the disadvantages of using language models (particularly neural network-based language models) as "black boxes" by allowing human users to understand the reasons behind the responses provided by the language model.

[0013] In one particular aspect, this specification describes a system implemented as a computer program on one or more computers at one or more locations that trains a language model to perform reasoning tasks.

[0014] The system includes a reward model configured to receive an input including a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query, and to process the input to compute a score indicating how successful the one or more reasoning steps were in producing a correct final answer to the query.

[0015] The system can train a reward model using multiple training examples using result supervised training or process supervised training. Typically, each training example includes a corresponding sample query text sequence that characterizes a corresponding sample query (i.e., a training query) and a corresponding reference response text sequence that includes a reference ("true value") final answer to the corresponding sample query. For each training example, the system can generate one or more candidate reasoning trajectories in response to the sample query of the training example. Each candidate reasoning trajectory includes one or more reasoning steps. The system can assign a target score to each reasoning step in each candidate reasoning trajectory based on the reference response text sequence in the training example, where the target score indicates whether the reasoning step is correct or incorrect. The system then trains the reward model to generate a reward score for the reasoning step in the candidate reasoning trajectory that matches the corresponding target score of the reasoning step.

[0016] In result supervised training, the system determines whether a candidate reasoning trajectory produces a reference final answer in a training example. If the candidate reasoning trajectory produces a reference final answer in a training example, the system assigns a reward score to each reasoning step in the candidate reasoning trajectory, indicating that the reasoning step is correct. If the candidate reasoning trajectory does not produce a final answer in a training example, the system assigns a reward score to each reasoning step in the candidate reasoning trajectory, indicating that the reasoning step is incorrect.

[0017] In process supervision training, the reference response text sequence includes a reference reasoning trajectory, which includes a sequence of reasoning steps with a final answer as the last reasoning step. The system assigns a target score to the current reasoning step based on whether the reasoning steps generated up to the current reasoning step match the sequence of reasoning steps in the reference reasoning trajectory. That is, the target score of the current reasoning step can be determined based on a comparison (e.g., a similarity score) between each reasoning step that has been generated up to the current reasoning step (including the current reasoning step) and the corresponding reasoning step in the sequence of reasoning steps in the reference reasoning trajectory.

[0018] To train the language model, the system may receive a training query and use the policy language model and the trained reward model to generate one or more expert reasoning trajectories in response to the training query, wherein each corresponding expert reasoning trajectory includes a corresponding plurality of reasoning steps. The system may use the one or more expert reasoning trajectories to train the language model. The policy language model may be a model configured to receive an input including a sequence of query text representing a query and one or more reasoning steps that have been generated in response to the query, and process the input to generate a next reasoning step.

[0019] In some embodiments, to generate an expert trajectory, the system generates a plurality of candidate expert reasoning trajectories in response to a training query using a policy language model. The system can generate a performance score for each candidate expert reasoning trajectory based on the reward scores generated for the reasoning steps in the candidate expert reasoning trajectories using a reward model. The system selects one or more candidate expert reasoning trajectories with the highest performance scores as the expert reasoning trajectories.

[0020] In some embodiments, to generate an expert trajectory, after one or more reasoning steps have been generated for the expert reasoning trajectory, the system uses a policy language model to generate multiple candidate next reasoning steps. The system uses a reward model to generate a reward score for each candidate next reasoning step and selects the candidate next reasoning step with the highest reward score as the next step of the expert reasoning trajectory. The system can repeat the above steps until the next step matches the final answer indicator (e.g., until the next reasoning step includes text corresponding to the final answer) or the maximum number of steps has been reached. The system can then use the resulting reasoning trajectory as the expert reasoning trajectory.

[0021] In another aspect, this specification describes a system, implemented as a computer program on one or more computers at one or more locations, that performs reasoning tasks using a language model.

[0022] The system may obtain an input text sequence representing an input query. The system also obtains a reward model that has been trained on a plurality of training examples, wherein each training example includes a corresponding sample query text sequence representing a sample query and a corresponding reference response text sequence, wherein the reference response text sequence includes at least a reference final answer to the sample query, and wherein the reward model is configured to process an input including a query text sequence representing the query and one or more reasoning steps that have been generated in response to the query to calculate a reward score indicating a degree of success of the one or more reasoning steps in producing a correct final answer to the query.

[0023] The system generates a plurality of candidate reasoning trajectories as responses to an input query using a language model. Each candidate reasoning trajectory includes a corresponding plurality of candidate reasoning steps that include a candidate final answer. The system may select a best reasoning trajectory from the plurality of candidate reasoning trajectories based on one or more reward scores calculated by a trained reward model, and output the best reasoning trajectory as an output response to the input query.

[0024] The present specification also describes a computer-implemented method executed by the above system. The present specification also describes one or more computer storage media storing instructions, which, when executed by one or more computers, cause one or more computers to execute the above method.

[0025] The subject matter described in this specification can be implemented in specific embodiments to realize one or more of the following advantages.

[0026] This specification provides a technique for using a reward model to guide reinforcement learning of a language model to perform reasoning tasks. Although the reward model can be trained and applied using purely result-based supervision or process-based supervision, the reward model can generate reward signals not only for the results of performing the task (e.g., the final answer to the query), but also for individual reasoning steps. By using the reward model to provide reward signals to the reasoning steps, the language model is improved when performing the reasoning task, for example, by reducing result errors and / or reasoning trajectory errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1A An example training system for training a language model to perform reasoning tasks is shown.

[0028] Figure 1B The reinforcement learning process performed by the training system is shown.

[0029] Figure 2A An example of using a reward model to select expert reasoning trajectories is shown.

[0030] Figure 2B An example of using a reward model to generate expert reasoning trajectories is shown.

[0031] Figure 3 is a flowchart of an example process for training a language model.

[0032] Figure 4 is a flowchart of an example process for performing an inference task using a trained language model.

[0033] Like reference numbers and names in the various drawings indicate like elements. DETAILED DESCRIPTION

[0034] This specification describes techniques for training a language model to perform reasoning tasks. Specifically, a language model is trained to generate a response in response to a query (e.g., a question), the response comprising a sequence of reasoning steps that lead to a final answer to the query (e.g., as the last reasoning step). The described techniques can be used to train a language model to perform a variety of reasoning tasks.

[0035] In one example, the trained language model can be used to diagnose faults in a mechanical system operating in a real-world environment. The input query can include observations about the mechanical system (e.g., from one or more sensors) and questions about the mechanical system, such as "Given the measurements from sensors A, B, and C, is the system working correctly?" The output response can provide the final answer to the question and the reasoning trajectory that led to the final answer.

[0036] In another example, the trained language model can be deployed as part of an educational software program that assists a user in learning or practicing one or more corresponding skills. For example, the input query can be a math word problem, and the output response includes the correct solution to the word problem and the intermediate steps to reach the correct solution. The output response can be provided to a student as a demonstration of how to solve a math problem, or used as a grading tool to provide feedback on a student's practice.

[0037] In another example, the trained language model can be used for natural language control of tasks in a real-world environment. The input query can be related to the task, for example, it can include a request to perform a task. The output response can be used to control, for example, a mechanical system (which can be referred to as a mechanical agent) or a computer system for performing a task. As an example, the input query can include, for example, a high-level question from a human to perform a task, such as "What is the most cost-effective way to fabricate this component with our current equipment?" The output response includes a final answer to a specified plan, and the reasoning steps for determining the plan. The output response can be used to control one or more mechanical agents, such as a robotic arm, to perform the plan. In some other examples, the reasoning trace in the output response can be used to provide a human-interpretable explanation, for example, the user can use the explanation to 1) determine whether to perform the action specified in the final answer in the output response or 2) later access to evaluate the performance of the language model when certain criteria are met or diagnose the cause of the error caused by the execution of the action specified in the response generated by the language model. In some implementations, the output responses may specify actions or routes to be taken by a human user to perform a task in the real-world environment, and the system may provide the output responses to the human user, such as via a display or audio output device, to guide the user in performing the task in the real-world environment.

[0038] Figure 1A An example training system 100 is shown for training a language model 110 to perform reasoning tasks. Training system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described below may be implemented.

[0039] Typically, language model 110 is trained to generate a response 130 in response to query 120 that includes a sequence of reasoning steps leading to a final answer to the query.

[0040] review Figure 1A , the language model 110 has a set of language model parameters 115. In some implementations, in each of a plurality of steps, the language model 110 may process an input including a query 120 to generate an output 150 that represents an inference step 130 in an inference trace responsive to the query 120. Each of the query 120 and the inference step 130 may include a sequence of text tokens.

[0041] The training system 100 includes a reinforcement learning engine 135 that performs reinforcement learning to update the language model parameters 115 of the language model 110 using a reward model 140. The reward model 140 is configured to process an input 150 that includes a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query 120 to calculate a reward score 160 that indicates how successful the one or more reasoning steps were in producing a correct final answer to the query.

[0042] The language model 110 may be a neural network configured to process an input to generate an output comprising a probability distribution over a set of text tokens in a text token vocabulary, wherein the probability of each token represents the likelihood that the text token immediately follows the input.

[0043] The text word-gram vocabulary may include any suitable word-grams that occur in natural language text, such as ASCII characters, words, word fragments, or n-grams of different distributions. For example, the text word-gram vocabulary may be fixed, or may have been generated by applying a suitable word-gram analyzer, such as a byte-pair encoding word-gram analyzer (tokenizer) or a SentencePiece word-gram analyzer, to a text corpus.

[0044] For example, language model 110 may be an autoregressive neural network. Language model 110 is referred to as an autoregressive neural network because the neural network autoregressively generates an output sequence of tokens by generating each particular token in the output sequence conditioned on a current input sequence, including any (e.g., all) tokens preceding a particular text token in the output sequence, i.e., tokens that have been generated for any previous positions preceding a particular position of a particular token in the output sequence, and contextual inputs that provide context for the output sequence (“context sequence”), such as a sequence of text tokens representing query 120.

[0045] For example, when generating a word-gram at any given position in the output sequence, the current input sequence may include a context sequence and word-grams at any previous positions before the given position in the output sequence. As a specific example, the current input sequence may include a context sequence followed by word-grams at any (e.g., all) previous positions before the given position in the output sequence. Optionally, the context and the current output sequence may be separated by one or more predetermined word-grams within the current input sequence.

[0046] More specifically, to generate a particular token at a particular position within the candidate output sequence, the neural network may process the current input sequence to generate a score distribution, such as a probability distribution, that assigns a corresponding score, such as a corresponding probability, to each text token in a vocabulary of text tokens. The neural network may then use the score distribution to select a text token from the vocabulary as the particular token. For example, the neural network may greedily select the highest-scoring token, or may sample tokens from the distribution, such as using kernel sampling or another sampling technique.

[0047] As a specific example, the language model 110 can be or include an autoregressive Transformer-based neural network that includes (i) a sequence of multiple attention blocks, each attention block applying a self-attention operation, and (ii) an output subnetwork that processes the output of the last attention block to generate a score distribution.

[0048] The neural network can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in: J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, DdL Casas, LA Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche,LAHendricks,M.Rauh,P.Huang,A.Glaese,J.Welbl,S.Dathathri,S.Huang,J.Uesato,J.Mellor,I.Higg ins,A.Creswell,N.McAleese,A.Wu,E.Elsen,SMJayakumar,E.Buchatskaya,D.Budden,E.Sutherland,K.Simonyan, M.Paganini,L.Sifre,L.Martens,XLLi,A.Kuncoro,A.Nemazadeh,E.Gribovskaya,D.Donato,A.Lazaridou,A.Mens ch,J.Lespiau,M.Tsimpoukelli,N.Grigorev,D.Fritz,T.Sottiaux,M.Pajarskas,T.Pohlen,Z.Gong,D.Toyama,C.de Masson d'Autume,Y.Li,T.Terzi,V.Mikulik,I.Babuschkin,A.Clark,D.de Las Casas,A.Guy,C.Jones,J.Bradbury,M.Johnson,BAHechtman,L.Weidinger,I.Gabriel,WSIsaac,E.Lockhart,S.Osindero,L.Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. .

[0049] However, in general, a Transformer-based neural network includes a sequence of attention blocks, and during processing of a given input sequence, each attention block in the sequence receives a corresponding input hidden state for each input token in the given input sequence. The attention block then updates the hidden state of the last token in the given input sequence at least in part by applying self-attention to generate a corresponding output hidden state for the last token. The input hidden state of the first attention block is the embedding of the input token in the input sequence, and the input hidden state of each subsequent attention block is the output hidden state generated by the previous attention block.

[0050] In this example, the output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last input token in the input sequence to generate a score distribution.

[0051] In some embodiments, prior to reinforcement learning, the language model 110 can be initialized using a base language model that has been pre-trained using unsupervised learning. That is, the language model 110 can start with the same architecture and weights as a base language model that has been trained on a text corpus to perform one or more language modeling tasks that do not require labeled training examples. For example, the system 100 or another training system can pre-train the base language model on the following: a masked language modeling (MLM) task for predicting masked portions of input text; a next sentence prediction (NSP) task for predicting whether the second sentence is a subsequent sentence of the first sentence given two sentences; and / or an autoregressive pre-training (ARP) task for predicting the next word in a sequence given a word-unit sequence. As a specific example, the language model 110 can be pre-trained with a maximum likelihood objective on a large text dataset (e.g., text publicly available from the Internet or another text corpus).

[0052] In some embodiments, prior to reinforcement learning and after initializing the language model 110, the system 100 can fine-tune the language model 110 using supervised fine-tuning (SFT) based on a set of supervised training examples. Each supervised example includes a training input word-gram sequence and a corresponding target output word-gram sequence. During SFT, the system 100 performs supervised learning to update the parameters of the language model 110, for example, to maximize the log-likelihood that the language model 110 outputs the target output word-gram sequence for the corresponding training input sequence. In some embodiments, the supervised training examples include examples of input queries and corresponding target reasoning trajectories for the input queries. That is, the training input word-gram sequence is a word-gram sequence representing the input query, and the target output word-gram sequence is a word-gram sequence representing the target reasoning trajectory.

[0053] In some of these embodiments, the system 100 may also perform contextual learning of the language model 110 by providing one or more hints in the input of the language model 110. Each hint may be an example of an input-output pair, where the input is an example of an input query and the output is an example of an output reasoning trace that should be generated in response to the input query.

[0054] The system 100 trains the reward model 140 (i.e., updates the parameters 145 of the reward model 140) using a set of training examples 148. Each training example 145 includes a corresponding query text sequence and a corresponding reference response text sequence that characterizes a corresponding sample query. The reference response text sequence includes at least a reference final answer to the corresponding sample query.

[0055] As described above, the reward model 140 is configured to receive an input 150 including a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query, and process the input 150 to calculate a reward score 160 indicating how successful the one or more reasoning steps are in producing a correct final answer to the query. In some embodiments, for each reasoning step in the reasoning trace generated in response to the query, the reward model 140 outputs a score for the reasoning step. In a specific example, the reward model 140 can be trained to generate a score for a reasoning step to indicate the likelihood that the reasoning step is "correct" or "incorrect."

[0056] In some embodiments, the reward model 140 is trained to generate a score for each reasoning step in the reasoning trace to indicate whether the reasoning trace containing the step leads to the correct final answer. The resulting reward model trained using this process is called an outcome supervised reward model (ORM) because its training is supervised by the reference final answer.

[0057] In some implementations of training the ORM, the system 100 obtains multiple candidate reasoning trajectories for each training example in response to a sample query of the training example. Figure 1BAs described, the system 100 can generate candidate reasoning trajectories using a policy language model. The system 100 assigns a target score to each reasoning step in each candidate reasoning trajectory based on a reference response text sequence in a training example. The target score indicates whether the reasoning step is correct or incorrect based on whether the corresponding final answer in the candidate reasoning trajectory matches the reference final answer of the training example. That is, if the candidate reasoning trajectory produces a reference final answer, the system 100 assigns a target score indicating that the reasoning step is correct to each reasoning step in the candidate reasoning trajectory. Otherwise, the system 100 assigns a target score indicating that the reasoning step is incorrect to each reasoning step in the candidate reasoning trajectory. Then, the system 100 trains the reward model 140 to generate a reward score 160 for the reasoning step in the candidate reasoning trajectory that matches the corresponding target score of the reasoning step. For example, the system 100 can train the reward model 140 using a supervised learning technique based on backpropagation.

[0058] In some embodiments, the reward model 140 is trained to generate a score for each specific reasoning step in a candidate reasoning trajectory to indicate whether the reasoning steps up to and including that step are correct. The resulting reward model trained using this process is called a process-supervised reward model (PRM) because its training is supervised by the process including the intermediate reasoning steps of the reasoning trajectory.

[0059] In some embodiments of training a PRM, the reference response text sequence for each training example includes a reference reasoning trajectory, which is a sequence of reasoning steps with a final answer as the last step. The system 100 assigns a target score to the current reasoning step based on whether the reasoning steps generated up to that point match the sequence of reasoning steps in the reference reasoning trajectory. That is, the target score for the current reasoning step can be determined based on a comparison (e.g., a similarity score) between each reasoning step that has been generated up to and including the current reasoning step and the corresponding reasoning step in the sequence of reasoning steps in the reference reasoning trajectory.

[0060] In some embodiments of training a PRM, the system 100 may use feedback from a human annotator to determine a target score for an inference step in a candidate inference trajectory. For example, the system 100 or another system may present to a human annotator (i) a sample query and reference response text sequence for a training example and (ii) a candidate inference trajectory obtained for the training example. The system 100 may then receive input from the annotator that identifies the first inference step in the candidate inference trajectory where a significant error (if any) exists. In one example, a significant error is defined as a step where the information expressed is incorrect, or where a correct solution is no longer possible without undoing the step. After receiving the input from the annotator, the system 100 may use the feedback from the annotator to assign a target score to the candidate inference step.

[0061] In some embodiments, the reward model 140 can be implemented as a language model that is configured to process an input sequence including a query and one or more reasoning steps and generate an output word element indicating the correctness of the reasoning step (e.g., "correct" or "incorrect"). The reward model 140 can also output the probability assigned to the "correct" word element by the language model as a reward score. In some embodiments, similar to the language model 110, the reward model 140 can be (i) initialized using a pre-trained basic language model, (ii) fine-tuned using SFT, and / or (iii) subjected to contextual learning through few-trial prompts.

[0062] Once the reward model 140 has been implemented, the reinforcement learning engine 135 performs reward model-based (RM-based) reinforcement learning to update the language model 110. During RM-based reinforcement learning, the reinforcement learning engine 130 updates a policy based on the reward score predicted by the reward model 140, which maps the inputs including (i) the query 120 and (ii) the reasoning steps 130 generated so far to the output representing the next reasoning step.

[0063] Figure 1B An example of a RM-based reinforcement learning process performed by the training system 100 is shown. In the example process described below, the system 100 uses an expert iteration method for reinforcement learning. Examples of implementing expert iteration include examples described in D. Silver, et al., "Mastering the game of go without human knowledge," Nature, 550(7676): 354-359, 2017 and T. Anthony, et al., "Thinking fast and slow with deep learning and tree search," Advances in Neural Information Processing Systems, 30, 2017. Typically, expert iteration alternates between two operations, which include (i) policy improvement and (ii) distillation.

[0064] exist Figure 1BIn the example shown, during policy refinement, system 100 uses policy language model 112 and trained reward model 140 to perform a search process to produce expert reasoning trace 190. During distillation, system 100 uses expert reasoning trace 190 to train language model 110, such as by using supervised learning.

[0065] The strategic language model 112 can be implemented using any suitable language model. In some embodiments, the strategic language model 112 can be the same model as the language model 110. That is, the language model 110 (e.g., after being subjected to pre-training, SFT, and / or contextual learning using few-trial prompts) can be used to generate the expert reasoning trajectory 190. The system 100 can perform multiple iterations of expert iterations. After the language model 110 has been trained in a particular iteration using the expert reasoning trajectory 190, the system 100 can use the language model 110 as the strategic language model 112 to generate the expert reasoning trajectory 190 for the next iteration. In some embodiments, the system can also use the strategic language model 112 to generate candidate reasoning trajectories for training the reward model 140.

[0066] To generate expert reasoning trace 190 , system 100 receives training query 170 and generates expert reasoning trace 190 in response to training query 170 using policy language model 112 and trained reward model 140 .

[0067] In some embodiments, the system can use the policy language model 112 to generate multiple candidate expert reasoning trajectories 180 (e.g., by random sampling) in response to the training query 170, generate a performance score 185 for each candidate expert reasoning trajectory using the trained reward model 140, and select one or more candidate expert reasoning trajectories with the highest performance scores as the expert reasoning trajectory 190.

[0068] Figure 2A An example of selecting an expert reasoning trajectory using a reward model 140 is shown. The system can determine a performance score 185 for a candidate expert reasoning trajectory based on a reward score generated by the reward model for each reasoning step in the candidate expert reasoning trajectory, for example, by summing the reward scores for corresponding reasoning steps in the candidate expert reasoning trajectory. The system selects the candidate expert reasoning trajectory with the highest performance score as the expert reasoning trajectory 190. In this case, the system can use a reward model 140 that has been trained using outcome supervised training, so the reward model (outcome supervised reward model, ORM) is configured to generate a reward score for each reasoning step to indicate whether the reasoning step is likely to have been in a reasoning trajectory that leads to the correct final answer. A strategy that maximizes the ORM score for each step typically maximizes the RM estimated probability that each step ultimately reaches the correct final answer.

[0069] Figure 2B An example of generating an expert reasoning trajectory 190 using a policy language model 112 and a reward model 140 is shown. In this case, after one or more reasoning steps have been generated for the expert reasoning trajectory, the system generates multiple candidate next reasoning steps 182 using the policy language model 112. The system generates a reward score (e.g., 185a, 185b, or 185c) for each candidate next reasoning step using the trained reward model. The system selects the candidate next reasoning step with the highest reward score as the next step of the expert reasoning trajectory, and repeats the process until the next step matches the final answer indicator (e.g., until the next reasoning step includes text corresponding to the final answer) or the maximum number of steps has been reached. In this case, the system can use a reward model 140 that has been trained using process supervision training, so the reward model (process supervision reward model, PRM) is configured to generate a reward score for the current reasoning step to indicate whether the reasoning steps up to the current reasoning step are correct. The strategy of maximizing the PRM score typically selects each step to maximize the RM estimated probability that the steps so far are correct. If the steps so far are correct, this typically means that such a strategy minimizes the probability of introducing an error in the current step.

[0070] Figure 3 1 is a flow chart of an example process 300 for training a language model to perform an inference task. For convenience, process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system appropriately programmed according to the present specification, such as training system 100 depicted in FIG. 1 , can perform process 300.

[0071] At 310, the system obtains a plurality of training examples. Each training example includes a corresponding sample query text sequence characterizing a corresponding sample query and a corresponding reference response text sequence including a reference final answer to the corresponding sample query. The system can obtain training examples from one or more of a variety of data sources, such as from a database or a labeled data set. As an illustrative example, the publicly available GSM8K data set includes math word problems and natural language solutions to the math word problems.

[0072] At 320, the system trains a reward model on a plurality of training examples. The reward model is configured to receive an input including a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query, and process the input to calculate a reward score indicating how successful the one or more reasoning steps are in producing a correct final answer to the query.

[0073] As described above, the system can train a reward model based on candidate reasoning trajectories and target scores assigned to reasoning steps in the candidate reasoning trajectories. The target scores can be assigned based on: (i) outcome-supervised training, where the training of the reward model is supervised by the reference final answers to the training examples, or (ii) process-supervised training, where the training of the reward model is supervised by the process represented by the reasoning trajectory.

[0074] At 330, the system uses the trained reward model to train the language model. As described above, the system performs RM-based reinforcement learning to update the parameters of the language model. That is, during reinforcement learning, the system uses the reward model to generate rewards. As described above, the reward model can be implemented as a language model that is configured to process an input sequence including a query and one or more reasoning steps, and generate output tokens indicating the correctness of the reasoning step (e.g., "correct" or "incorrect"). The reward model is used to provide reward signals to individual reasoning steps during reinforcement learning. This is in contrast to conventional techniques, in which language models are rewarded only based on the correctness of the final answer during reinforcement learning. As will be discussed with reference to Figures 5A-6, by using a reward model to provide reward signals to reasoning steps, the language model is improved when performing reasoning tasks, such as by reducing result errors and / or reasoning trajectory errors.

[0075] Figure 4 1-3. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations.

[0076] The system receives an input query at 410. The input query can be represented by a sequence of text word-grams.

[0077] At 420, the system processes the input query using the trained language model to generate a plurality of candidate output reasoning traces in response to the input query. Each candidate output reasoning trace is a sequence of text tokens representing a sequence of reasoning steps, wherein the last reasoning step in the reasoning trace is or includes a final answer generated in response to the input query. As described with reference to FIG. 1 , since the language model is autoregressive and is configured to process the current input sequence to generate a probability distribution of tokens in a vocabulary for the next token in the output, the system can generate multiple different candidate output sequences in response to the same query using the same model by sampling from the probability distribution.

[0078] At 430, the system selects the best reasoning trajectory from the candidate output reasoning trajectories using a ranking strategy. In some embodiments, the system may select the best reasoning trajectory to be output using a reward model that has been trained using the training process described with reference to FIG. 1. Specifically, the system may use the trained reward model to calculate the performance score of each candidate output reasoning trajectory. In some embodiments, the system may select the candidate output reasoning trajectory with the highest performance score as the output reasoning trajectory. In some other embodiments, the system may weight the final answer in each candidate output reasoning trajectory with the corresponding performance score, and identify the final answer with the maximum total weight as the "correct" final answer. That is, the total weight of each final answer can be determined by summing the performance scores of each output reasoning trajectory with the final answer, and then the "correct" final answer is identified as the final answer with the maximum total weight. The system can then select among the candidate output reasoning trajectories that lead to the identified "correct" final answer based on the performance score.

[0079] At 440, the system outputs the best reasoning trajectory. In some embodiments, to ensure the quality of the generated response, the system may choose not to provide an output response when the performance score of the best reasoning trajectory estimated by the reward model is below a threshold. That is, before outputting the best reasoning trajectory as an output response, the system may determine whether the score of the best reasoning trajectory estimated by the reward model is below a threshold, and only output the best reasoning trajectory as an output response in response to determining that the score is not below the threshold.

[0080] This specification uses the term "configured" in connection with system and computer program components. For a system of one or more computers to be configured to perform specific operations or actions, it means that the system has installed thereon software, firmware, hardware, or a combination thereof that, when operated, causes the system to perform those operations or actions. For one or more computer programs to be configured to perform specific operations or actions, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform those operations or actions.

[0081] The embodiments and functional operations of the subject matter described in this specification may be implemented in digital electronic circuits, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, for example, one or more computer program instruction modules encoded on a tangible non-transitory storage medium, for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device.

[0082] The term "data processing apparatus" refers to data processing hardware and covers all types of apparatus, devices and machines for processing data, including, for example, a programmable processor, a computer or multiple processors or computers. The apparatus may also be or further include special-purpose logic circuits, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the device may optionally include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0083] A computer program, which may also be referred to or described as a program, software, software application, application, module, software module, script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to a related program, or in multiple coordinated files, such as files storing one or more modules, subroutines, or portions of code. A computer program may be deployed and executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communications network.

[0084] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or at all, and may be stored on storage devices in one or more locations. Thus, for example, an index database may include multiple collections of data, each of which may be organized and accessed in a different way.

[0085] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines can be installed and run on the same one or more computers.

[0086] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by a dedicated logic circuit such as an FPGA or ASIC, or a combination of a dedicated logic circuit and one or more programmed computers.

[0087] Computers suitable for executing computer programs can be based on general or special microprocessors or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented or incorporated therein by a dedicated logic circuit. Typically, the computer will also include or be operably coupled to receive data from one or more large-capacity storage devices (such as disks, magneto-optical disks or optical disks) for storing data or to transmit data or both. However, the computer does not need to have such a device. In addition, the computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.

[0088] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks.

[0089] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with a user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and the input from the user may be received in any form, including acoustic, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's device in response to a request received from the web browser. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user in return.

[0090] The data processing apparatus for implementing the machine learning model may also include, for example, dedicated hardware accelerator units for processing common and computationally intensive parts of machine learning training or production, such as inference, workloads.

[0091] Machine learning models can be implemented and deployed using a machine learning framework such as the TensorFlow framework.

[0092] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component, such as a data server, or includes a middleware component, such as an application server, or includes a front-end component, such as a client computer with a graphical user interface, a web browser, or an application, through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by digital data communication of any form or medium, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0093] A computing system may include a client and a server. The client and the server are usually remote from each other and typically interact through a communication network. The relationship between the client and the server is by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data such as HTML pages to a user device, for example, for displaying data to a user interacting with a device acting as a client and receiving user input from the user. Data generated at the user device, such as the result of a user interaction, can be received from the device at the server.

[0094] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or the scope that can be claimed, but should be interpreted as descriptions of the features of specific embodiments of specific inventions. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations and even initially claimed to be so protected, in some cases, one or more features from the claimed combination may be cut out from the combination, and the claimed combination may be directed to a sub-combination or a variant of the sub-combination.

[0095] Similarly, although operations are depicted in a particular order in the figures and described in the claims, this should not be construed as requiring that such operations be performed in the particular order or sequential order shown, or that all of the operations shown be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0096] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method for training a language model to perform an inference task, the method comprising: Obtaining a plurality of training examples, wherein each training example includes a corresponding sample query text sequence representing a corresponding sample query and a corresponding reference response text sequence including a reference final answer to the corresponding sample query; training a reward model on the plurality of training examples, wherein the reward model is configured to receive an input comprising a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query, and process the input to compute a reward score indicating how successful the one or more reasoning steps are in producing a correct final answer to the query; and The language model is trained using the trained reward model.

2. The method according to claim 1, wherein: Training the language model using the trained reward model includes: receiving training inquiries; generating one or more expert reasoning traces in response to the training query using the policy language model and the trained reward model, each respective expert reasoning trace comprising a respective plurality of reasoning steps; and The language model is trained using the one or more expert reasoning traces.

3. The method according to claim 1 or claim 2, wherein: Training the reward model on the plurality of training examples includes, for each training example in one or more of the training examples: In response to the sample query of the training example, a plurality of candidate reasoning trajectories are obtained, wherein each candidate reasoning trajectory includes one or more reasoning steps; Assigning a target score to each reasoning step in each candidate reasoning trajectory based on the reference response text sequence in the training example, wherein the target score indicates whether the reasoning step is correct or incorrect; and The reward model is trained to generate reward scores for reasoning steps in the candidate reasoning trajectories that match corresponding target scores for the reasoning steps.

4. The method according to claim 3, wherein: Assigning the target score to each reasoning step in the candidate reasoning trajectory includes: determining whether the candidate reasoning trajectory produces the reference final answer in the training example; In response to determining that the candidate reasoning trajectory produces the reference final answer in the training example, assigning a target score to each reasoning step in the candidate reasoning trajectory indicating that the reasoning step is correct; and In response to determining that the candidate reasoning trajectory does not produce the final answer in the training example, each reasoning step in the candidate reasoning trajectory is assigned a target score indicating that the reasoning step is incorrect.

5. The method according to claim 3, wherein: For each training example, the reference response text sequence includes a reference reasoning trajectory including a sequence of reasoning steps, the sequence of reasoning steps including the final answer as the last reasoning step, and wherein assigning the target score to each reasoning step in the candidate reasoning trajectory includes: The target score is assigned to the current reasoning step based on whether each reasoning step that has been generated up to the current reasoning step matches the sequence of reasoning steps in the reference reasoning trajectory.

6. A method according to any one of claims 3 to 5 and also dependent on claim 2, wherein: The strategic language model is configured to receive input including a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query, and process the input to generate a next reasoning step.

7. The method according to claim 6, wherein: Generating the expert trajectory includes: generating a plurality of candidate expert reasoning trajectories in response to the training query using the policy language model; generating, using the reward model, a performance score for each of the candidate expert reasoning traces based on: a corresponding reward score generated for each of one or more of the reasoning steps in the candidate expert reasoning trace; and One or more candidate expert reasoning trajectories having the highest performance scores are selected as the expert reasoning trajectories.

8. The method according to claim 6, wherein: Generating an expert trajectory in the expert trajectory trajectory includes: After one or more reasoning steps have been generated for the expert reasoning trajectory, generating a plurality of candidate next reasoning steps using the policy language model; generating a reward score for each of the candidate next reasoning steps using the reward model; Selecting the candidate next reasoning step with the highest reward score as the next step of the expert reasoning trajectory; and The using and selecting steps are repeated until the next step matches the final answer indicator or a maximum number of steps has been reached.

9. The method according to any preceding claim, further comprising: Get the basic language model.

10. The method according to claim 8 and also dependent on claim 2, further comprising: The basic language model is updated to generate the policy language model.

11. The method according to claim 10, wherein: Updating the basic language model to generate the policy model includes: Supervised fine-tuning is performed on the base language model on the supervised training examples.

12. The method according to any one of claims 9 to 11, wherein: Training the language model includes: The language model is initialized based on the base language model.

13. The method according to any one of claims 9 to 12, wherein: Training the reward model includes: The reward model is initialized based on the base language model.

14. The method according to any preceding claim, further comprising: The reward model is updated using the trained language model.

15. The method of any preceding claim further comprising using the trained language model to generate an output response to an input query related to a real-world environment, the output response providing information about the real-world environment or specifying an action or route to be taken in the real-world environment.

16. The method according to claim 15, wherein: The input query is related to a task in the real-world environment, and the method further includes using the generated output response to control one or more mechanical agents or computer systems acting in the real-world environment to perform the task.

17. The method according to claim 15, wherein: The input query is related to a task in the real-world environment, and the output response specifies an action or route to be taken in the real-world environment. The method also includes providing the output response to a user for guiding the user to perform the task in the real-world environment.

18. The method of any one of claims 1 to 15, further comprising using the trained language model to generate an output response to an input query, the input query comprising observations about a mechanical system operating in a real-world environment, the generated output response being used to diagnose a fault in the mechanical system.

19. A computer-implemented method for performing an inference task using a language model, the method comprising: obtaining an input text sequence representing an input query; Obtaining a reward model that has been trained on a plurality of training examples, wherein each training example comprises a corresponding sample query text sequence representing a sample query and a corresponding reference response text sequence, wherein the reference response text sequence comprises at least a reference final answer to the sample query, and wherein the reward model is configured to process an input comprising a query text sequence representing a query and one or more reasoning steps that have been generated in response to the query to calculate a reward score indicating a degree of success of the one or more reasoning steps in producing a correct final answer to the query; generating a plurality of candidate reasoning trajectories as responses to the input query using the language model, each candidate reasoning trajectory comprising a corresponding plurality of candidate reasoning steps comprising a candidate final answer; selecting a best reasoning trajectory from the plurality of candidate reasoning trajectories based on one or more reward scores calculated by the trained reward model; and The optimal inference trajectory is output as an output response to the input query.

20. The method according to claim 19, wherein: Selecting the best inference trajectory from the plurality of candidate inference trajectories based on one or more scores calculated by the trained reward model includes: generating a corresponding weight for each of the plurality of candidate reasoning trajectories using the reward model, wherein the corresponding weight measures an estimated probability of correctness of the corresponding candidate reasoning trajectory; selecting the best final answer as the candidate final answer having the largest total weight; and Among the candidate reasoning trajectories that produce the optimal final answer, a candidate reasoning trajectory having a highest weight estimated by the reward model is selected as the best reasoning trajectory.

21. The method of claim 19 or claim 20, wherein: The language model has been trained using a method according to any one of claims 1 to 18.

22. The method according to any one of claims 19 to 21, wherein: The input text sequence specifies a math word problem, and the output response specifies a solution to the math problem and reasoning steps for solving the math word problem.

23. The method according to any one of claims 19 to 22, further comprising: Before outputting the best inference trajectory as the output response, determining whether a score estimated by the reward model for the best inference trajectory is below a threshold; as well as The best inference trajectory is output as the output response only in response to determining that the score is not below the threshold.

24. The method according to any one of claims 19 to 23, wherein: The input query relates to a real-world environment, and the output response provides information about the real-world environment or specifies an action or route to be taken in the real-world environment.

25. The method according to claim 24, wherein: The input query is related to a task in a real-world environment, and the method further includes using the output response to control one or more mechanical agents or computer systems in the real-world environment to perform the task.

26. The method according to claim 24, wherein: The input query is related to a task in the real-world environment, and the output response specifies an action or route to be taken in the real-world environment. The method also includes providing the output response to a user for guiding the user to perform the task in the real-world environment.

27. The method according to any one of claims 19 to 24, wherein: The input queries include observations about a mechanical system operating in a real-world environment, and the output responses are used to diagnose faults in the mechanical system.

28. A method according to any preceding claim, wherein: The language model includes a neural network.

29. A system comprising: one or more computers; as well as One or more storage devices storing instructions which, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 28.

30. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method of any one of claims 1 to 28.

Citation Information

Cited By

  • Enhanced feedback-based medical interactive large model training method and system

    CN120809166A