Reinforcement Learning with Option Feedback

By using a pairwise preference function to compare outputs from a target and alternative neural networks and updating weights based on preference scores, the system effectively trains neural networks to prioritize high-quality outputs, addressing limitations in existing methods and reducing reward hacking susceptibility.

JP7698093B1Active Publication Date: 2025-06-24ジーディーエム·ホールディング·エルエルシー
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024058938
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-02-14
Filing Date
2024-04-01
Publication Date
2025-06-24
Estimated Expiration
2044-04-01

AI Technical Summary

Technical Problem

Existing machine learning model training methods often struggle to accurately prioritize high-quality outputs, especially in conditional generation tasks, due to limitations in modeling preferences and susceptibility to reward hacking.

Method used

The system employs a pairwise preference function to train a target neural network by comparing network outputs from the target and alternative neural networks, using a preference score to update weights and encourage prioritization of higher-quality outputs, while also regularizing with a reference neural network to mitigate reward hacking.

Benefits of technology

This approach enables more accurate and efficient training of neural networks based on preferences, achieving a specific performance threshold in conditional generation tasks with fewer iterations and reduced susceptibility to reward hacking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698093000001_ABST
    Figure 0007698093000001_ABST
Patent Text Reader

Abstract

Provided are a system device, a method, and a program for training a neural network by comparing the quality between network outputs. 【Solution means】The method includes: receiving one or more network inputs; for each of the network inputs, processing the network input using a target neural network to generate a first network output, and processing the network input using an alternative neural network to generate a second network output; applying a preference function to the first network output and the second network output to generate a preference score for comparing the first network output and the second network output; and updating the weights of the target neural network using an objective function that encourages the first network output to be preferred over the second network output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to processing data using a machine learning model.

Background Art

[0002] A machine learning model receives an input and generates an output, such as a predicted output, based on the received input. Some machine learning models are parametric models and generate an output based on the received input and the values of the model's parameters.

[0003] Some machine learning models are deep models that use multiple model layers to generate an output for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to the received input to generate an output.

Summary of the Invention

Problems to be Solved by the Invention

[0004] This specification describes a system and method implemented as a computer program on one or more computers in one or more locations that can train a neural network using a preference function that compares a measure of quality between network outputs.

Means for Solving the Problems

[0005] According to a first aspect, there is provided a method for training a target neural network, which is executed by one or more computers, has weights of a plurality of target neural networks, and is configured to process network inputs according to the weights of the target neural network to generate network outputs. In each of a plurality of training steps, the method includes receiving one or more network inputs. For each of the network inputs, the method includes processing the network input using the target neural network to generate a first network output, processing the network input using an alternative neural network for the training step to generate a second network output, and applying a preference function to the first network output and the second network output to generate a preference score indicating the likelihood that the first network output is a higher quality output than the second network output. For each training step, the method includes updating the weights of the target neural network using an objective function that includes a first term that encourages prioritizing the first network output over the corresponding second network output according to the preference score.

[0006] In some implementations, the method further includes, in each of the plurality of training steps and for each of the one or more network inputs, (i) determining a first likelihood score assigned to the first network output by the target neural network when the network input is provided, and (ii) determining a second likelihood score assigned to the first network output by a reference neural network for the training step when the network input is provided, and the objective function includes a second term that penalizes the target neural network to generate a first likelihood score that deviates from the corresponding second likelihood score.

[0007] In some implementations, in each of a plurality of training steps, the reference neural network for the training step is the alternative neural network for the training step.

[0008] In some implementations, to generate a preference score indicating the likelihood that a first network output is a higher-quality output than a second network output, the step of applying a preference function to the first network output and the second network output comprises: (i) processing the first network output and the second network output using a preference model neural network to generate a preference network output, wherein the preference model neural network is trained to generate a corresponding preference network output for each processed pair of neural network outputs, the processed pair of neural network outputs indicating the likelihood that one member of the processed pair is a higher-quality output than the other member of the processed pair; and (ii) generating a preference score indicating the likelihood that the first network output is a higher-quality output than the second network output based on the preference network output.

[0009] In some implementations, in each of a plurality of training steps, and for each of one or more network inputs, when the network input is provided, the step of determining a second likelihood score assigned to a first network output by a reference neural network for the training step comprises determining the second likelihood score assigned to the first network output based on a geometric mixture of a first likelihood score of the first network output and a likelihood score of the first network output under a fixed reference distribution, and based on a normalization constant for the training step.

[0010] In some implementations, in each of a plurality of training steps, an alternative neural network for a training step has weights of a plurality of alternative neural networks for the training step, has the same architecture as a target neural network, and the weights of the alternative neural network for the training step are determined by an exponentially weighted moving average of the weights of the target neural network over the current and previous training steps.

[0011] In some implementations, the objective function is based on a preference score determined by a preference function indicating the likelihood that a first network output is a higher quality output than a second network output, for each network input and corresponding first and second network outputs, a first likelihood score of the first network output determined by the target neural network, and a second likelihood score of the first network output determined by a fixed reference distribution.

[0012] In some implementations, the step of updating the weights of the current target neural network using an objective function that includes a first term that encourages preferring a first network input over a corresponding second network input according to the preference score includes determining a gradient of the objective function with respect to the weights of the target neural network and updating the weights of the target neural network based on the gradient of the objective function.

[0013] In some implementations, the step of determining the gradient of the objective function with respect to the weights of the current target neural network includes, for each network input and corresponding first and second network outputs, the gradient of the first likelihood score of the first network output determined by the target neural network with respect to the weights of the current target neural network, the preference score determined by a preference function indicating the likelihood that the first network output is a higher-quality output than the second network output, the first likelihood score of the first network output determined by the target neural network, and the likelihood score of the first network output determined by a fixed reference distribution.

[0014] In some implementations, the target neural network is a large language model.

[0015] In some implementations, the preference neural network is a large language model.

[0016] In some implementations, the first network output is a first network output sequence, the second network output is a second network output sequence, and the step of processing the network input using an alternative neural network to generate the second network output includes the step of autoregressively generating the second network output sequence using the alternative neural network.

[0017] In some implementations, in each of a plurality of training steps, and for each of one or more network inputs, when a network input is provided, the step of determining a second likelihood score assigned to a first network output by a reference neural network for the training step includes, for each output element of the first network output sequence, when a network input is provided, using a target neural network to determine a first likelihood score for the output element, and determining a second likelihood score for the output element based on a geometric mixture of the first likelihood score of the output element and the likelihood score of the output element under a fixed reference distribution, and based on a normalization constant for the training step, and determining a second likelihood score assigned to the first network output based on the second likelihood scores of the output elements within the first network output sequence.

[0018] In some implementations, the preference function can model non-transitive preference relations.

[0019] According to another aspect, there is provided a system comprising one or more computers and one or more storage devices communicatively coupled to the one or more computers, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method described above.

[0020] According to another aspect, there are provided one or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method described above.

[0021] Particular embodiments of the subject matter described herein can be implemented to realize one or more of the following advantages.

[0022] The described system can use a pairwise preference function to train a target neural network that processes pairs of network outputs to determine preference scores. In some implementations, the described system can train a preference model to model the pairwise preference function. Unlike conventional methods for training neural networks that often relied on fitting a reward model to assign numerical values to individual network outputs according to preference, the described system can model non-transitive preferences. By using a pairwise preference function, it is also possible to model transitive preferences more accurately than conventional methods. Thus, the described system can accurately train neural networks based on preferences for a wider range of applications (e.g., modeling human preferences, predicting the outcome of a combat game, etc.) than conventional methods.

[0023] By using a pairwise preference function to more accurately model the preferences of a conditional generation task, the described system can train the target neural network to achieve a specific threshold of performance (e.g., determined based on a preference function) in the conditional generation task with fewer training iterations than conventional methods. Thus, the described system can train neural networks based on preferences more efficiently (e.g., with respect to training time, computational resources, etc.) than conventional methods.

[0024] Conventional methods often train neural networks to optimize a reward model adapted based on preference data, which may be vulnerable to the influence of reward hacking. In reward hacking, network outputs that maximize rewards under the reward model are not prioritized over network outputs that receive worse rewards. In contrast, the described system can train a target neural network to approach a Nash equilibrium of preferences with respect to a set of alternative neural networks (e.g., to generate network outputs that are at least as preferable as the network outputs generated by the set of alternative neural networks).

[0025] In particular, the described system can train the target neural network by comparing the output from the target neural network with the corresponding output of an alternative neural network. The described system can use an objective function that includes a term indicating the preference of the target neural network output compared to the corresponding alternative neural network output. While training the target neural network, the described system can update the alternative neural network at each training step based on the training of the target neural network. By comparing the preference of the target neural network output with the corresponding network output from the alternative neural network updated at each training step, the objective function of the described system can encourage the target neural network to approach a Nash equilibrium of preferences with respect to the set of alternative neural networks. By training the target neural network to approach a Nash equilibrium of preferences, the described system can be less susceptible to the influence of reward hacking compared to conventional methods.

[0026] The described system can also regularize the training of the target neural network with respect to the reference neural network. In particular, the described system can use an objective function that includes a term that penalizes the target neural network for being too different from the reference neural network. For example, the reference neural network can produce a known safe or desirable output, and the described system can train the target neural network based on preferences while producing an output similar to that of the reference neural network.

[0027] Accordingly, the described system can enable more accurate and efficient training of the neural network based on preferences while ensuring that the target neural network can produce a preferred output.

[0028] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0029]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0030] Like reference numerals and designations in the various drawings refer to like elements.

[0031] FIG. 1 shows an exemplary training system 100. The training system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations where the systems, components, and techniques described below are implemented.

[0032] The training system 100 can train a target neural network 102 using a set of training data 104 to perform a conditional generation task. In particular, the system 100 can train the target neural network 102 to perform a conditional generation task tailored to the preferences of the task. The target neural network 102 has, for example, a set of weights of the target neural network including the weights of the layers of the target neural network and optionally biases. The target neural network 102 can be a neural network trained (e.g., pre-trained) to perform a conditional generation task, and the system 100 can further train (e.g., fine-tune) the target neural network 102 to perform a conditional generation task tailored to the preferences of the task.

[0033] The training system 100 includes a weight update system 106 that can update the weights of the target neural network 102 to train the target neural network 102 over a series of training steps. In each training step, the training system 100 can train the target neural network 102 to perform a conditional generation task on a set of network inputs 108 (e.g., a mini-batch).

[0034] The target neural network 102 can be any generative model suitable for generating output samples based on the network input 108. The target neural network 102 will be described in more detail below with reference to FIG. 2.

[0035] The conditional generation task can be any of a variety of tasks. As an example, the network input 108 can be a text prompt, and the conditional generation task can generate an output sample as described by the text prompt. As a further example, the target neural network 102 can be configured to output, for example, a sample image, a sample video, a sample audio, etc., as described by a text prompt from the network input 108.

[0036] The conditional generation task can be a style transfer task for matching a style such as an image, a video, an audio, etc., specified by the network input 108. For example, the target neural network 102 can be configured to output, for example, a sample image, a sample video, a sample audio, etc., that matches the style of the network input 108.

[0037] The conditional generation task can be a transcription task in which the target neural network 102 is configured to generate a transcription of text such as a video, an audio, etc., specified by the network input 108.

[0038] The conditional generation task can be a summarization task in which the target neural network 102 is configured to generate a text summary of, for example, text, an image, a video, an audio, etc., specified by the network input 108.

[0039] The conditional generation task can be a natural language processing (NLP) task, and the target neural network 102 can be a large language model configured to perform a language processing task. For example, the target neural network 102 can be configured to generate a text response to a network input 108 provided by a user. As another example, the target neural network 102 can be configured to generate code for a computer program (e.g., in a programming language such as Python, Java, C++) to perform a task specified by a network input 108 provided by a user.

[0040] The conditional generation task can be an action selection task for an agent that interacts with an environment. For example, the network input 108 can be, for example, the agent's sensor data that characterizes an observation of the environment, the observed trajectories of other agents in the environment, etc., and the target neural network 102 can be configured to sample an action of the agent based on the observation of the environment. As a further example, the network input 108 can be, for example, LIDAR data, camera data, etc., collected for an autonomous vehicle, and the target neural network 102 can be configured to select an action for controlling the autonomous vehicle. As another example, the network input 108 can characterize the state of a game, and the target neural network 102 can be configured to select an action of a player of the game.

[0041] The training system 100 trains the target neural network 102 to perform a conditional generation task using the pairwise preference function 109. The pairwise preference function 109 can compare a given output sample from the target neural network 102 with a corresponding output sample from the alternative neural network 110 and generate a preference score 111 indicating the likelihood that the output sample from the target neural network 102 is of higher quality than the output sample from the alternative neural network 110.

[0042] The alternative neural network 110 can be any of a variety of networks that can provide alternative output samples (for the target neural network 102) for the conditional generation task. Examples of the alternative neural network 110 will be described below with reference to FIG. 5.

[0043] As an example, the preference score 111 can indicate which of the output sample from the target neural network 102 or the corresponding output sample from the alternative neural network 110 is more likely to be drawn from the ground truth distribution.

[0044] As another example, for the conditional generation task, the preference score 111 can indicate which of the output sample from the target neural network 102 or the corresponding output sample from the alternative neural network 110 is more likely to have, for example, less distortion, perceptual distortion, etc.

[0045] As another example, the preference score 111 can indicate which of the output sample from the target neural network 102 or the corresponding output sample from the alternative neural network 110 is more likely to receive a greater reward for the conditional generation task.

[0046] As yet another example, the preference score 111 can quantify human preference between an output sample from the target neural network 102 and a corresponding output sample from the alternative neural network 110.

[0047] In particular, the preference function 109 can model non - transitive preferences between network outputs. That is, for network outputs A, B, and C, when the preference function determines that A is more preferred than B (i.e., is more likely to be a high - quality output) and that B is more preferred than C, the preference function 109 can determine that C is more preferred than A. In the above example, in a model limited to modeling transitive preferences, necessarily, A would be determined to be more preferred than C. Non - transitive preferences are seen in various applications (e.g., modeling human preferences, predicting player performance in a combat game, etc.). Thus, the preference function 109 can more accurately model preferences for a wider range of applications compared to models limited to modeling transitive preferences.

[0048] In each training step, the system 100 uses the pairwise preference function 109 to compare a target network output 112 that includes an output sample from the target neural network 102 with a corresponding alternative network output 114 that includes an output sample from the alternative neural network 110, to generate a preference score 111 for the training step. Based on the preference score 111, the weight update system 106 generates an update 116 of the weights of the target neural network 102 and an update 118 of the alternative neural network 110. The weight update system 106 can generate the weight updates 116 and 118 to train the target neural network 102 to perform a conditional generation task. The process of training the target neural network 102 will be described in more detail below with reference to FIG. 3.

[0049] FIG. 2 shows an exemplary target neural network 102.

[0050] The target neural network 102 includes a generative model 202 that can process the network input 108 and generate an output sample 204 to perform a conditional generation task.

[0051] The generative model 202 can have any suitable architecture for processing the network input 108 and generating the output sample 204 to perform a conditional generation task. As an example, the generative model 202 can be an autoregressive generative model (e.g., Transformer, recurrent neural network, etc.) that can autoregressively generate a sequence as the output sample 204 based on the network input 108. The generative model can be, for example, a large language model (LLM) that can autoregressively generate a tokenized representation of text data, a vision - language model (VLM) that can autoregressively generate a tokenized representation of image or video data, an audio - language model that can autoregressively generate a tokenized representation of text data, etc.

[0052] As another example, the generative model 202 can be a diffusion model (e.g., denoising diffusion model, score - based diffusion model, latent diffusion model, etc.) that can generate the output sample 204 by repeatedly transforming samples from a noise distribution (e.g., Gaussian distribution) based on the network input 108 over a series of iterations. For example, the generative model 202 can be a diffusion model that uses a denoising neural network having any suitable architecture (e.g., convolutional neural network, recurrent neural network, etc.) to transform samples from a noise distribution.

[0053] As another example, the generative model 202 can be a neural network that generates the output sample 204 by transforming samples from a noise distribution (e.g., a Gaussian distribution). The generative model 202 can be, for example, the generator network of a generative adversarial network, the decoder of a variational autoencoder, a normalizing flow, and the like.

[0054] The generative model 202 can generate a network output 208 that defines the respective conditional probability or likelihood of the output sample 204 when the corresponding network input 108 is provided. For example, the network output 208 can include the conditional probability for a set of values for each of the output samples 204 when the corresponding network input is provided. As another example, the network output 208 can determine statistical characteristics (e.g., mean, variance, etc.) regarding the conditional distribution for each of the output samples 204 when the corresponding network input 108 is provided. As another example, the generative model 202 can include a plurality of processing layers and a plurality of layer activation functions, and the network output 208 can include output values from the processing layers and layer activation functions resulting from the generative model 202 processing the corresponding network input 108.

[0055] The target neural network 102 can include a sample likelihood system 206. The sample likelihood system 206 can process data specifying the conditional distribution 208 to generate a likelihood score 210 that specifies the likelihood of each of the output samples 204 under the conditional distribution 208 when the corresponding network input 108 is provided.

[0056] The likelihood score 210 can be any of a variety of numerical values indicating the conditional probability of the likelihood of the output sample 204 when the corresponding network input 108 is provided. For example, the likelihood score 210 can be the conditional probability or likelihood of the output sample 204 when the corresponding network input 108 is provided. As another example, the likelihood score 210 can be a transformation of the conditional probability or likelihood of the output sample 204 (e.g., log probability, log likelihood, negative log likelihood, non-normalized value, etc.) when the network input 108 is provided. As another example, the sample likelihood system 206 can determine an auxiliary distribution that approximates the conditional distribution of the output sample 204 when the corresponding network input 108 is provided, and the likelihood score 210 can be the probability or likelihood of the output sample 204 under the auxiliary distribution.

[0057] The sample likelihood system 206 can determine a likelihood score 210 based on the network output 208 in any of a variety of ways. For example, the network output 208 can include a set of conditional probabilities of the output sample 204, and the sample likelihood system can return the conditional probabilities of the output sample 204 as the likelihood score 210. As another example, the network output 208 can include data specifying a conditional distribution of the output sample 204, and the sample likelihood system 206 can return the likelihood of the output sample 204 under the specified conditional distribution as the likelihood score 210. As another example, the network output 208 can include data specifying a conditional distribution of the output sample 204, and the sample likelihood system 206 can determine an approximate likelihood of the output sample 204 under the specified conditional distribution and return it as the likelihood score 210. As a specific example, the network output 208 can include data specifying a layer output from the generative model 202, and the sample likelihood system can determine an approximate likelihood (e.g., by determining an approximation of the conditional distribution using importance weights) to return as the likelihood score 210. As another specific example, the sample likelihood system 206 can include a neural network trained to approximate the conditional distribution of the output sample 204 when given the network input 108, and can process the network input 108 to determine an approximate likelihood of the output sample 204 to return as the likelihood score 210.

[0058] FIG. 3 is a flowchart of an exemplary process for training a target neural network to perform a conditional generation task. For convenience, process 300 is described as being performed by a system of one or more computers located in one or more locations. For example, a training system such as training system 100 of FIG. 1 appropriately programmed in accordance with this specification can execute process 300.

[0059] The training system can train a target neural network over a plurality of training steps.

[0060] In each training step, the training system can receive a set of one or more network inputs for the training step (step 302). For example, the system can receive a mini-batch of network inputs from a set of training data.

[0061] The system can process the network inputs for the training step using the target neural network to generate a corresponding network output (step 304). In particular, the target neural network can process each of the network inputs to generate a corresponding output sample for performing a conditional generation task.

[0062] The system can process the network inputs for the training step using an alternative neural network for the training step to generate a corresponding network output (step 306). In particular, the alternative neural network can process each of the network inputs to generate a corresponding output sample for performing a conditional generation task.

[0063] For each of the network inputs, the system can apply a preference function to the corresponding network outputs from the target neural network and the alternative neural network to determine a preference score for comparing the network outputs (step 308). For each pair of network outputs from the target neural network and the alternative neural network, the preference score indicates the likelihood that the network output from the target neural network is of higher quality than the network output from the alternative neural network.

[0064] In some implementations, the preference function is a preference model neural network. The preference model neural network can be trained to process pairs of neural network outputs and generate corresponding preference network outputs for each processed pair of neural network outputs that indicate the likelihood that one member of the processed pair has a higher quality output than the other member of the processed pair. An exemplary process for training the preference model neural network will be described in more detail below with reference to FIG. 4.

[0065] The system can determine whether the training step is complete (step 310). For example, the system can determine that the training step is complete when it has processed all of the network inputs within a set of network inputs. If the training step is not complete, the system can process the next network input for the training step.

[0066] When the training step is complete, the system can update the weights of the target neural network based on the preference scores determined for the training step (step 312). In particular, the system can use an objective function for the training step that includes terms that encourage prioritizing the network outputs from the target neural network over the corresponding network outputs from the alternative neural network according to the preference scores to update the weights of the target neural network. An exemplary objective function for the training step will be described in more detail below with reference to FIG. 5.

[0067] FIG. 4 is a flow diagram of an exemplary process for training a preference model neural network. For convenience, process 400 is described as being performed by a system of one or more computers located in one or more locations. For example, a training system such as training system 100 of FIG. 1 properly programmed in accordance with this specification can perform process 400.

[0068] The system receives a set of training data (step 402). The training data includes exemplary network inputs for training a target neural network to perform a conditional generation task. The training data can include data specifying a preference function for network outputs. For example, the training data can include, for each exemplary network input, a pair of exemplary network outputs, and a corresponding indication that one network output of the pair is of higher quality than the other network output. As a specific example, the training data can include, for each pair of exemplary network outputs, a target likelihood that the network output of the pair is of higher quality than the other output of the pair. As another specific example, the training data can include, for each pair of exemplary network outputs, a target classification of which network output of the pair is the higher quality output.

[0069] The system trains a preference model neural network (step 404) to model a preference function specified by training data. As described above, the system processes pairs of neural network outputs and generates corresponding preference network outputs for each processed pair of neural network outputs that indicate the likelihood that one member of the processed pair has a higher quality output than the other member of the processed pair. The preference model neural network can be trained. As an example, the system can train the preference model neural network to replicate target likelihoods from a training data set by optimizing a regression loss function (e.g., L2 loss) between the preference network output and the corresponding target likelihood. As another example, the system can train the preference model neural network to replicate target classifications for pairs of network outputs by optimizing a classification loss function (e.g., cross-entropy loss) between the preference network output and the corresponding target classification.

[0070] The preference model neural network can have any of a variety of network architectures suitable for processing pairs of network outputs. For example, if the network output is an image, the preference model neural network can include a neural network suitable for processing image data (e.g., a convolutional neural network, a vision transformer, etc.). As another example, if the network output is a sequence of values (e.g., time series data, text sequence, video sequence, audio sequence, etc.), the preference model neural network can include a neural network suitable for processing sequences (e.g., a recurrent neural network, a transformer network, etc.). As a specific example, the preference model neural network can be a large language model.

[0071] The system can train a preference model neural network to model any of various preference functions of network outputs. For example, for each pair of network outputs, the preference function can indicate the likelihood that one network output of the pair is of higher quality than the other network output, as determined by some quality measure. As a further example, the preference function can indicate the likelihood that one output of the pair has lower, for example, distortion, perceptual distortion, etc. As yet another further example, when the network output is an action of an agent that interacts with the environment, the preference function can indicate the likelihood that one output of the pair receives a greater reward.

[0072] The preference function can represent the preferences of a human or a group of humans. For example, for each pair of network outputs, the preference function can indicate the human preference for one of the network outputs as the higher quality output of the pair. As another example, the preference function can indicate the human belief or conviction that one of the network outputs is more likely to be the higher quality output of the pair.

[0073] In particular, the preference function can be a non - transitive preference function. For example, if y1, y2, and y3 are network outputs, the preference function can also indicate that a human prefers y1 over y2, y2 over y3, but also y3 over y1. As another example, when the network output represents an action performed by an agent interacting within an environment, the preference function can indicate the non - transitive likelihood of success of the output action when executed by competing agents.

[0074] As described above, the system can then use the training dataset and the trained preference model neural network to train a target neural network to perform a conditional generation task (step 406).

[0075] FIG. 5 is a flow diagram of an exemplary process for determining an objective function for training a target neural network to perform a conditional generation task. For convenience, process 500 is described as being performed by a system of one or more computers located in one or more locations. For example, a training system such as training system 100 of FIG. 1 appropriately programmed in accordance with this specification can perform process 500.

[0076] As described above, the system trains the target neural network over a series of training steps. In each training step, the system receives a set of network inputs.

[0077] The system processes the set of network inputs to generate a set of network outputs from the target neural network and a corresponding set of network outputs from an alternative neural network for the training step (step 502).

[0078] The alternative neural network for each time step can be determined based on the target neural network. As described above, the alternative neural network can be any of a variety of networks that can provide alternative output samples (for the target neural network) of the conditional generation task.

[0079] For example, an alternative neural network can be a network trained to perform a conditional generation task with a different network architecture compared to the target neural network. As another example, the alternative neural network can have the same network architecture as the target neural network. As a further example, the alternative neural network can have the same architecture as the target neural network but with different hyperparameters (e.g., trained with different hyperparameters, performing inference using different hyperparameters, etc.), such as network size, network depth, number of inputs, learning rate, etc. As a further example, the alternative neural network can be a reparameterization of the target neural network with different network weights, or can be a copy of the target neural network fine-tuned on a specific set of training data for the alternative neural network. As a further example, the alternative neural network can process modified network inputs compared to the target neural network (e.g., by conditionally generating network outputs based on modified neural network inputs that specify generating outputs in a different style than the target neural network, by processing a modified input prompt compared to the target neural network, etc.).

[0080] In particular, the alternative neural network can have the same network architecture as the target neural network and can have weights of the alternative neural network determined based on the weights of the target neural network in the current and previous training steps. As an example, the weights of the alternative neural network can be the weights of the target neural network from the previous time step (e.g., the weights of the alternative neural network at the T-th training step,

[0081]

Number

[0082] is θ T-K and can be the weights of the target neural network at the T-K-th time step).

[0083] As another example, the weights of the alternative neural network can be determined by the moving average of the weights of the target neural network at the current and previous training steps. That is, the alternative neural network models the conditional distribution

[0084]

Number

[0085] and the weights of the alternative neural network at the T-th training step

[0086]

Number

[0087] is determined by the moving average.

[0088]

Number

[0089] K > 0 is the window length, and θ i are the weights of the target neural network at the i-th training step.

[0090] As another example, the weights of the alternative neural network can be determined by the exponentially weighted moving average of the weights of the target neural network at the current and previous training steps. That is, the alternative neural network models the conditional distribution

[0091] [Number]

[0092] can be modeled, and the weights of the alternative neural network at the T-th training step

[0093] [Number]

[0094] are determined by an exponential moving average.

[0095] [Number]

[0096] α ∈ [0, 1] is a weighting parameter, and θ i is the weight of the target neural network at the i-th training step.

[0097] The system uses a pairwise preference function to process each network output from the target neural network and the corresponding network output from the alternative neural network at the time step to determine a pairwise preference score (step 504).

[0098] In some implementations, the system can determine a likelihood score for each network output according to the target neural network (step 506). The likelihood score for a particular network output by the target neural network indicates the likelihood of the network output under the conditional distribution modeled by the target neural network when a network input corresponding to the particular network output is given. For example, if the target neural network has a conditional distribution p θWhen sampling the network output y based on the network input x using θ, θ is the weight of the target neural network (e.g., y ~ p θ (·|x)), and the likelihood score of y according to the target neural network can be as follows. p θ (y|x)

[0099] In some implementations, the network output can be an output sequence. For example, for a particular network output, y can be a sequence of N elements y = (y1,..., y N ). The target neural network can generate the output sequence autoregressively, and the likelihood score by the target neural network for the output y = (y1,..., y N ) can be as follows.

[0100] [Number]

[0101] Wherein, p θ (y i |x, y j<i ) is the likelihood score of the i-th element y j<i by the target neural network when given the network input x and the preceding elements y i of the output sequence.

[0102] In some implementations, the system can determine the likelihood score of each network output according to a reference neural network for the training step (step 508). The reference neural network can be any generative model suitable for sampling the network output conditioned on the network input.

[0103] For example, the reference neural network can be a network trained to perform a conditional generation task with a different network architecture compared to the target neural network. As another example, the reference neural network can have the same network architecture as the target neural network. As a further example, the reference neural network can have the same architecture as the target neural network using different hyperparameters (e.g., trained with different hyperparameters, performing inference using different hyperparameters, etc.), such as network size, network depth, number of inputs, learning rate, etc. As a further example, the reference neural network can be a reparameterization of the target neural network with different network weights, or can be a copy of the target neural network fine-tuned on a specific set of training data for the reference neural network. As a further example, the reference neural network can process a modified network input compared to the target neural network (e.g., by generating a network output conditionally based on a modified neural network input that specifies generating an output in a different style than the target neural network, by processing a modified input prompt compared to the target neural network, etc.).

[0104] The likelihood score for a particular network output by the reference neural network at a time step indicates the likelihood of the network output under the conditional distribution modeled by the reference neural network given the network input corresponding to the particular network output. For example, if q represents the conditional distribution modeled by the reference neural network, the likelihood score for a network output y given a network input x according to the reference neural network can be as follows. q(y|x)

[0105] The reference neural network for each time step can be determined based on the target neural network. In particular, the reference neural network can have the same network architecture as the target neural network and can have the weights of an alternative neural network determined based on the weights of the target neural network in the current and previous training steps. As an example, the weights of the reference neural network can be the weights of the target neural network from the previous time step (e.g., the weights of the reference neural network in the T-th training step,

[0106]

Number

[0107] can be θ T-K and be the weights of the target neural network in the (T - K)-th time step).

[0108] As another example, the weights of the reference neural network can be determined by the moving average of the weights of the target neural network in the current and previous training steps. That is, the reference neural network can model the conditional distribution

[0109]

Number

[0110] and the weights of the reference neural network in the T-th training step

[0111]

Number

[0112] is determined by the moving average.

[0113]

Number

[0114] K > 0 is the window length, and θ i is the weight of the target neural network at the i-th training step.

[0115] As another example, the weights of the reference neural network can be determined by the exponentially weighted moving average of the weights of the target neural network at the current and previous training steps. That is, the reference neural network models the conditional distribution

[0116]

Number

[0117] and the weights

[0118]

Number

[0119] of the reference neural network at the T-th training step are determined by the exponentially weighted moving average.

[0120]

Number

[0121] α ∈ [0, 1] is the weighting parameter, and θ i is the weight of the target neural network at the i-th training step.

[0122] The reference neural network can model a distribution determined by a fixed reference distribution μ(y|x). When the reference neural network is determined by the fixed reference distribution μ(y|x), the system can use the reference neural network to regularize the training of the target neural network. For example, the fixed reference distribution can be a previously trained neural network that has been verified to generate safe or desirable network outputs.

[0123] The reference neural network can model a distribution based on a mixture of the target neural network and the fixed reference distribution μ(y|x).

[0124] As an example, the reference neural network can be determined as an additive mixture between the target neural network and the fixed reference distribution μ(y|x). For example, using a weighting parameter A ∈ [0,1], the reference neural network is determined as follows. q(y|x)=Ap θ (y|x)+(1 - A)μ(y|x)

[0125] As another example, the reference neural network can be determined as a geometric mixture between the target neural network and the fixed reference distribution μ(y|x). For example, using a weighting parameter β ∈ [0,1], the reference neural network can be determined as follows.

[0126]

Number

[0127] Where c(x) is a normalization constant based on the network input x. In particular, c(x) can be defined as follows.

[0128]

Number

[0129] When the network output is the output sequence, the reference neural network can autoregressively process the output sequence, and the likelihood score by the reference neural network for the output y=(y1,…,y N ) can be as follows.

[0130]

Equation

[0131] where q(y i |x,y j<i ) is the likelihood score of the i-th element y j<i by the reference neural network when the network input x and the preceding elements y i of the output sequence are given.

[0132] For example, using the weighting parameter β∈[0,1], when the reference neural network is a geometric distribution between the target neural network and the fixed reference distribution μ(x|y), the likelihood score of the i-th element y i by the reference neural network can be determined as follows.

[0133]

Equation

[0134] where c(x,y j<i ) is the normalization constant based on the network input x and the preceding elements y j<i . In particular, c(x) can be defined as follows.

[0135]

Equation

[0136] In some implementations, the reference neural network is an alternative neural network.

[0137] The system then updates the target neural network using an objective function to perform a conditional sampling task (step 510).

[0138] The objective function includes a first term that encourages the network output from the target neural network to be preferred over the corresponding network output from the alternative neural network according to a preference score.

[0139] In some implementations, the objective function can include a second term that penalizes the target neural network for producing a network output having a likelihood score by the target neural network that deviates from the corresponding likelihood score by the reference neural network.

[0140] For example, using a weighting parameter η ∈ [0,1], the objective function can be as follows.

[0141]

Equation

[0142] where x is the network input, p θ is the conditional distribution modeled by the target neural network, p' is the conditional distribution modeled by the alternative neural network, and q is the conditional distribution modeled by the reference neural network.

[0143]

Equation

[0144] The preference score determined for the target network output y with respect to the alternative network output y', and KL is the Kullback-Leibler information measure.

[0145] The system can determine an objective function as the sum of the loss terms for each network input x. As an example, the system can determine the loss terms for the target neural network output y and the alternative neural network output y' when the network input x is given as follows.

[0146]

Number

[0147] Loss term

[0148]

Number

[0149] encourages the target neural network to produce an output that is preferred over the output generated by the alternative neural network. The loss term

[0150]

Number

[0151] penalizes the target neural network to generate a network output having a likelihood score by the target neural network that deviates from the corresponding likelihood score by the reference neural network.

[0152] In some implementations, the system can update the target neural network by calculating the gradient of the objective function with respect to the weights θ of the target neural network. For example, the system can determine the gradient of the objective function as the sum of the gradients of the loss terms for each network input, as follows.

[0153]

Number

[0154] As a specific example, the preference score

[0155]

Number

[0156] is a value between 0 and 1 indicating the likelihood that y is a higher-quality output than y', and when the reference neural network is a geometric mixture between the target neural network with weighting parameter β and the fixed reference distribution μ(y|x), the system can determine the gradient of the loss term for each network input as follows.

[0157]

Number

[0158] The system can update the weights of the alternative and reference neural networks (step 512). The system can determine whether to update the weights of the alternative and reference neural networks based on any of a variety of criteria. For example, the system can update the weights of the alternative and reference neural networks for each training step.

[0159] In particular, the system can update the weights of the alternative and reference neural networks using the updated weights of the target neural network for the training step. For example, if the alternative neural network is determined based on the average of the weights of the target neural network over a plurality of time steps, the system can update the weights of the alternative neural network to include the updated weights of the target neural network. As another example, if the reference neural network is determined based on the average of the weights of the target neural network over a plurality of time steps, the system can update the weights of the reference neural network to include the updated weights of the target neural network. As another example, if the reference neural network is determined based on a mixed distribution between the target neural network and a fixed reference distribution, the system can update the mixed distribution of the reference neural network to include the updated target neural network.

[0160] This specification uses the term "configured" with respect to system and computer program components. To be configured to perform a particular operation or action by one or more computer systems means that the system has installed in the computer, during operation, software, firmware, hardware, or combinations thereof that cause the system to perform the operation or action. To be configured to perform a particular operation or action by one or more computer programs means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.

[0161] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, tangibly implemented computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or one or more combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver apparatus for execution by a data processing apparatus.

[0162] The term “data processing apparatus” refers to data processing hardware and includes any kind of apparatus, device, and machine for processing data, e.g., a programmable processor, a computer, or multiple processors or computers. The apparatus can be, or further include, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0163] A computer program, also called a program, software, software application, app, module, software module, script, or code, or sometimes described as such, can be written in any form of programming language, including compiled or interpreted languages, declarative languages, or procedural languages, and it can be deployed in any form, such as as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program does not necessarily have to, but can correspond to a file in a file system. The program can be stored in a file dedicated to the program, in a part of a file that holds other programs or data, such as one or more scripts stored in a markup language document, or in multiple coordinated files, such as files that store one or more modules, subprograms, or parts of the code. A computer program can be deployed to be executed on one computer, or located at one site, or distributed across multiple sites and executed on multiple computers interconnected by a data communication network.

[0164] As used herein, the term "engine" is widely used to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and executed on the same one or more computers.

[0165] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, a special-purpose logic circuit such as an FPGA or ASIC, or by a combination of special-purpose logic circuits and one or more programmed computers.

[0166] Computers suitable for the execution of a computer program can be based on general purpose or special purpose microprocessors, or both, or any other kind of central processing unit. In general, a central processing unit receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuits. In general, a computer also includes one or more mass storage devices for storing data, such as, by way of example only, magnetic, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, transfer data to, or both, one or more mass storage devices. However, a computer need not have such devices. Further, a computer can be embedded in another device, such as, by way of example only a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.

[0167] Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices such as, for example, EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and all forms of non-volatile memory, media, and memory devices including CD-ROM and DVD-ROM disks.

[0168] To provide for interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user, and a keyboard and a pointing device such as, for example, a mouse or a trackball by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well, and, for example, feedback provided to the user can be any form of sensory feedback such as, for example, visual feedback, auditory feedback, or tactile feedback, and input received from the user can be received in any form including acoustic, speech, or tactile input. Further, the computer can interact with the user by sending and receiving documents between the devices used by the user, such as, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device such as, for example, a smartphone running a messaging application and receiving, in turn, a response message from the user.

[0169] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing common portions and computationally intensive portions of machine learning training or making, i.e., inference, workloads.

[0170] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework or the Jax framework.

[0171] Embodiments of the subject matter described herein can be implemented in a computing system that includes, for example, backend components as a data server, or middleware components such as an application server, or frontend components such as a graphical user interface, a web browser, or a client computer having an app through which a user can interact with an implementation of the subject matter described herein, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0172] The computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between a client and a server arises from computer programs that run on respective computers and have a client-server relationship to each other. In some embodiments, the server sends data, such as an HTML page, to a user device to display data to a user interacting with a device operating as a client and to receive user input from the user. For example, data generated at a user device, such as as a result of user interaction, can be received at the server from the device.

[0173] This specification includes many specific implementation details, but these are not limitations on the scope of any invention or what may be claimed, but rather should be construed as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features are described above as acting in certain combinations and were initially claimed as such, but in some cases, one or more features from the claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.

[0174] Similarly, operations are shown in the drawings and described in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or sequentially, or that all of the operations shown be performed to achieve a desired result. In some situations, multitasking and parallel processing may be advantageous. Further, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be incorporated together into a single software product or packaged into multiple software products.

[0175] Particular embodiments of the subject matter are described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve a desired result. As an example, the processes shown in the accompanying drawings do not necessarily require the particular or sequential order shown to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous.

Explanation of Symbols

[0176] 100 Training System 102 Target Neural Network 106 Weight Update System 108 Network Input 109 Pairwise Preference Function 110 Alternative Neural Network 111 Preference Score 112 Target Network Output 114 Alternative Network Output 116 Weight Update 202 Generation Model 204 Output Sample 206 Sample Likelihood System 208 Network Output 210 Likelihood Score 300 Process 400 Process 500 Process

Claims

1. 1. A method, executed by one or more computers, for training a target neural network, the target neural network having a plurality of target neural network weights and configured to process network inputs according to the target neural network weights to generate network outputs, wherein in each of a plurality of training steps: receiving one or more network inputs; For each of the one or more network inputs: processing the network inputs with the target neural network to generate a first network output; processing the network inputs using an alternative neural network for the training step to generate a second network output; applying a preference function to the first network output and the second network output to generate a preference score indicating the likelihood that the first network output is a higher quality output than the second network output; updating weights of the target neural network using an objective function including a first term that encourages a preference for the first network output over a corresponding second network output according to the preference score; A method comprising:

2. In each of the plurality of training steps, and for each of the one or more network inputs, determining a first likelihood score to be assigned by the target neural network to the first network output given the network input; determining a second likelihood score to be assigned to the first network output by a reference neural network for the training step given the network input; Further comprising: the objective function includes a second term that penalizes the target neural network for producing a first likelihood score that deviates from a corresponding second likelihood score. The method of claim 1.

3. 3. The method of claim 2, wherein in each of the plurality of training steps, the reference neural network for the training step is the alternative neural network for the training step.

4. applying the preference function to the first network output and the second network output to generate the preference score indicative of the likelihood that the first network output is a higher quality output than the second network output; processing the first network output and the second network output using a preference model neural network to generate preference network outputs, the preference model neural network being trained to process pairs of neural network outputs and to generate a corresponding preference network output for each processed pair of neural network outputs that indicates a likelihood that one member of the processed pair is a higher quality output than the other member of the processed pair; generating the preference score based on the preference network outputs, the preference score indicating the likelihood that the first network output is a higher quality output than the second network output; 2. The method of claim 1, comprising:

5. In each of the plurality of training steps, and for each of the one or more network inputs, determining the second likelihood score to be assigned to the first network output by a reference neural network for the training step given the network input, comprising: determining the second likelihood score to be assigned to the first network output based on a geometric mixture of the first likelihood score of the first network output and a likelihood score of the first network output under a fixed reference distribution and based on a normalization constant for the training step.

3. The method of claim 2, comprising:

6. In each of the plurality of training steps, the alternative neural network for the training step includes a plurality of alternative neural network weights for the training step and has the same architecture as the target neural network; the weights of the alternative neural network for the training step are determined by an exponential moving average of the weights of the target neural network over the current and previous training steps; The method of claim 1.

7. The objective function is, for each network input and corresponding first and second network outputs: a preference score determined by the preference function indicating the likelihood that the first network output is a higher quality output than the second network output; a first likelihood score of the first network output determined by the target neural network; and a second likelihood score of the first network output determined by a fixed reference distribution; The method of claim 1, based on

8. updating weights of the current target neural network using the objective function including a first term that encourages a preference for a first network input over a corresponding second network input according to the preference score; determining a gradient of the objective function with respect to the weights of the target neural network; updating weights of the target neural network based on the gradient of the objective function; 2. The method of claim 1, comprising:

9. The step of determining the gradient of the objective function with respect to the current target neural network weights comprises, for each network input and corresponding first and second network outputs: a gradient of a first likelihood score of the first network output as determined by the target neural network with respect to the current target neural network weights; a preference score determined by the preference function indicating the likelihood that the first network output is a higher quality output than the second network output; the first likelihood score of the first network output determined by the target neural network; and a likelihood score of the first network output, determined by a fixed reference distribution; and The method of claim 8, comprising the step of calculating:

10. The method of claim 1 , wherein the target neural network is a large language model.

11. The method of claim 4 , wherein the preference model neural network is a large language model.

12. the first network output comprises a first network output sequence; the second network output comprises a second network output sequence; processing the network inputs using the alternative neural network to generate the second network outputs comprises autoregressively generating the second network output sequence using the alternative neural network. The method of claim 2.

13. In each of the plurality of training steps, and for each of the one or more network inputs, determining the second likelihood score to be assigned to the first network output by a reference neural network for the training step given the network input, comprising: For each output element of the first network output sequence: determining a first likelihood score for the output elements using the target neural network given the network inputs; determining a second likelihood score for the output element based on a geometric mixture of the first likelihood score for the output element and a likelihood score for the output element under a fixed reference distribution and based on a normalization constant for the training step; determining the second likelihood score assigned to the first network output based on the second likelihood scores of the output elements in the first network output sequence; 13. The method of claim 12, comprising:

14. The method of claim 1 , wherein the preference function is capable of modeling a non-transitive preference relationship.

15. One or more computers; and one or more storage devices communicatively coupled to the one or more computers, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for training a target neural network, the target neural network having a plurality of target neural network weights, the target neural network being configured to process network inputs in accordance with the target neural network weights to generate network outputs, the operations comprising, in each of a plurality of training steps: Receiving one or more network inputs; For each of the one or more network inputs: processing the network inputs with the target neural network to generate a first network output; processing the network inputs using an alternative neural network for the training step to generate a second network output; applying a preference function to the first network output and the second network output to generate a preference score indicating the likelihood that the first network output is a higher quality output than the second network output; updating weights of the target neural network using an objective function including a first term that encourages a preference for the first network output over a corresponding second network output according to the preference score. system.

16. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations for training a target neural network, the target neural network having a plurality of target neural network weights and configured to process network inputs according to the target neural network weights to generate network outputs, the operations comprising, in each of a plurality of training steps: Receiving one or more network inputs; For each of the one or more network inputs: processing the network inputs with the target neural network to generate a first network output; processing the network inputs using an alternative neural network for the training step to generate a second network output; applying a preference function to the first network output and the second network output to generate a preference score indicating the likelihood that the first network output is a higher quality output than the second network output; updating weights of the target neural network using an objective function including a first term that encourages a preference for the first network output over a corresponding second network output according to the preference score; [0036] one or more non-transitory computer storage media,

17. said operation comprising: in each of said plurality of training steps, and for each of said one or more network inputs: determining a first likelihood score to be assigned by the target neural network to the first network output given the network input; determining a second likelihood score to be assigned to the first network output by a reference neural network for the training step given the network input; the objective function includes a second term that penalizes the target neural network for producing a first likelihood score that deviates from a corresponding second likelihood score.

17. One or more non-transitory computer storage media as recited in claim 16.

18. applying the preference function to the first network output and the second network output to generate the preference score indicating the likelihood that the first network output is a higher quality output than the second network output; processing the first network output and the second network output using a preference model neural network to generate preference network outputs, the preference model neural network being trained to process pairs of neural network outputs and generate a corresponding preference network output for each processed pair of neural network outputs that indicates a likelihood that one member of the processed pair is a higher quality output than the other member of the processed pair; generating the preference score based on the preference network output, the preference score indicating the likelihood that the first network output is a higher quality output than the second network output; 20. The one or more non-transitory computer storage media of claim 16, comprising:

19. The objective function is, for each network input and corresponding first and second network outputs: a preference score determined by the preference function indicating the likelihood that the first network output is a higher quality output than the second network output; a first likelihood score of the first network output determined by the target neural network; and a second likelihood score of the first network output determined by a fixed reference distribution; 17. The one or more non-transitory computer storage media of claim 16,

20. 17. The one or more non-transitory computer storage media of claim 16, wherein the preference functions are capable of modeling non-transitive preference relationships.

Citation Information

Patent Citations

  • Method and system for enhancing anti-attack capability of model based on adversarial samples

    CN111046394A

  • Method and system for testing robustness of artificial intelligence model

    CN112766315A

  • Neural network with improved performance, using automatically uncovered failure cases

    JP2023109726A

  • Inference device, inference method, and program

    JP2023534518A

  • Constrained Reinforcement Learning Neural Network System Using Pareto Front Optimization

    JP2023545021A