Training a sequence generation neural network using mass fraction

By using the optimal completion distillation technology and quality scores in neural network training, the problems of insufficient training performance and excessive computing resource consumption in the existing technology are solved, and high-performance neural network training is achieved.

CN111727442BActive Publication Date: 2025-06-24GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980013555.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-23
Filing Date
2019-05-23
Publication Date
2025-06-24
Estimated Expiration
2039-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high performance and excessive computing resource consumption when training sequences generate neural networks.

Method used

Optimal completion distillation (OCD) technology is used to train neural networks through quality scores, reduce dependence on computing resources, and improve the performance of sequence generation tasks.

Benefits of technology

Neural network training with advanced performance in sequence generation tasks is realized, without hyperparameter adjustment or pre-training, reducing the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111727442B_ABST
    Figure CN111727442B_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus for training a sequence generation neural network, including a computer program encoded on a computer storage medium. One of the methods includes: obtaining a batch of training examples; for each of the training examples: using a neural network to process the training network input in the training example to generate an output sequence; for each specific output position in the output sequence: identifying the head word of the system output at a position before the specific output position included in the output sequence; for each possible system output in the vocabulary, determining the highest quality score that can be assigned to any candidate output sequence including the head word followed by the possible system output; and determining an update to the current values of the network parameters, the update increasing the likelihood that the neural network generates a system output having a high quality score at that position.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] This specification relates to training neural networks.

[0002] A neural network is a machine learning model that uses one or more layers of non-linear units to predict an output for received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network - that is, the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of a corresponding set of parameters.

[0003] Some neural networks are recurrent neural networks. A recurrent neural network is a neural network that receives a sequence of inputs and generates a sequence of outputs from the input sequence. In particular, a recurrent neural network can use some or all of the internal state of the network from a previous time step when computing the output at the current time step. SUMMARY OF THE INVENTION

[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations that trains a sequence generation neural network to generate an output sequence conditioned on a network input.

[0005] In particular, the system uses "optimal completion distillation" to train the neural network. In optimal completion distillation, the system uses a sequence generation neural network to generate an output sequence, and then uses a quality score to train the neural network, where the quality score measures the quality of a candidate output sequence determined using the heads of words within the generated output sequence relative to a ground truth output sequence that the neural network should have generated. This is contrary to conventional techniques such as maximum likelihood estimation (MLE), in which the heads of words from the ground truth output sequence are directly provided as input to the neural network.

[0006] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.

[0007] By using quality scores in the training of a neural network as described in this specification, a system can train the neural network to have advanced performance on sequence generation tasks such as speech recognition or another sequence generation task. In fact, the described optimal completion distillation technique has no hyperparameters, and the neural network does not require any pre-training to achieve this level of performance, thus reducing the amount of computational resources consumed by the overall training process. Additionally, by efficiently identifying the highest score for each position in the output sequence as described in this specification, the amount of computational resources required to perform the training is reduced. Therefore, the described techniques allow the neural network to be trained to have advanced performance without excessive consumption of computational resources. As a specific example of the performance that a neural network trained using the described techniques can achieve, Table 1 (illustrated in Figure 6 shows the performance of training the same neural network on the same task: the Wall Street Journal speech recognition task using three different techniques - OCD (the technique described in this specification), scheduled sampling (SS), and MLE. In particular, Table 1 shows the performance in terms of word error rate (a lower word error rate indicates better performance) of neural networks trained using these three methods with various beam sizes. The beam size refers to the size of the beam used in beam search decoding during inference, i.e., after training, to generate the output sequence. As can be seen from Table 1, the performance of the neural network when trained using OCD far exceeds the performance of using the other two techniques, which were previously considered advanced network training techniques.

[0008] Details of one or more embodiments of the subject matter described in this specification are set forth in the following drawings and description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 Shows an example neural network training system.

[0010] Figure 2 Is a flowchart of an example process for training a sequence generation neural network.

[0011] Figure 3 Is a flowchart of an example process for generating Q values for a given position in an output sequence.

[0012] Figure 4 Shows an example of applying the OCD training technique when the quality metric is based on edit distance.

[0013] Figure 5 Is a flowchart of an example process for determining the highest quality score of a specific system output preceded by a specific word head when the quality metric is based on edit distance.

[0014] Figure 6 Shows the performance differences in training a neural network using three different techniques.

[0015] In the various figures, like reference numerals and names indicate like elements. Detailed Description

[0016] This specification describes a system implemented as a computer program on one or more computers in one or more locations that trains a neural network to generate an output sequence conditioned on a network input.

[0017] The neural network can be configured to generate any one of various output sequences conditioned on any one of various network inputs.

[0018] For example, the neural network can be a machine translation neural network. That is, if the network input is a sequence of words in an original language, such as a sentence or phrase, then the output sequence can be a translation of the input sequence into a target language, i.e., a sequence of words in the target language that represents the sequence of words in the original language.

[0019] As another example, the neural network can be a speech recognition neural network. That is, if the network input is a sequence of audio data representing a spoken utterance, then the output sequence can be a sequence of letters, characters, or words representing the utterance, i.e., a transcription of the input sequence.

[0020] As another example, the neural network can be a natural language processing neural network. For example, if the network input is a sequence of words in an original language, such as a sentence or phrase, then the output sequence can be a summary of the input sequence in the original language, i.e., a sequence with fewer words than the input sequence but retaining the essential meaning of the input sequence. As another example, if the network input is a sequence of words forming a question, then the output sequence can be a sequence of words forming an answer to the question.

[0021] As another example, the neural network can be part of a computer-aided medical diagnosis system. For example, the network input can be data from an electronic medical record (which can include physiological measurements in some examples) and the output sequence can be a sequence of predicted therapies and / or medical diagnoses.

[0022] As another example, the neural network can be part of an image processing system. For example, the network input can be an image and the output can be a sequence of text describing the image. As another example, the network input can be a sequence of text or different contexts and the output sequence can be an image describing the context. As another example, the network input can be image, audio, and / or video data and the output can be a sequence defining an enhanced version of the data (e.g., with reduced noise).

[0023] A neural network can have any of a variety of architectures. For example, a neural network can have an encoder neural network for encoding network inputs and a decoder neural network for generating an output sequence from the encoded network inputs. In some examples, the decoder is an autoregressive neural network, such as a recurrent neural network or an autoregressive convolutional neural network or an autoregressive attention-based neural network.

[0024] Figure 1 An exemplary neural network training system 100 is shown. Neural network training system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, in which the following systems, components, and techniques are implemented.

[0025] Neural network training system 100 trains a sequence generation neural network 110 with parameters 112 (referred to in this specification as "network parameters") that generates an output sequence. As described above, sequence generation neural network 110 can be configured to generate any of a variety of output sequences conditioned on any of a variety of network inputs.

[0026] In particular, sequence generation neural network 110 includes a decoder neural network that generates an output sequence conditioned on a system input - that is, directly or through a representation of the system input generated by an encoder neural network - step by step in time. At each time step, a recurrent neural network generates a likelihood distribution conditioned on previous outputs in the output sequence and the system input and over possible system outputs in the vocabulary of the system output, i.e., a score distribution including respective scores for each possible system output in the vocabulary. System 100 then selects the output at the time step by sampling from the likelihood distribution or selecting the most highly scored possible system output.

[0027] The sequence generation neural network can generally be any kind of neural network that generates an output defining a respective likelihood distribution over possible system outputs for each time step in the output sequence. Examples of such types of neural networks include sequence-to-sequence recurrent neural networks, self-attention-based neural networks, and convolutional neural networks.

[0028] System 100 trains sequence generation neural network 110 on training data to determine a trained value of network parameters 112 according to an initial value of network parameters 112 using an iterative training process.

[0029] Training data typically includes a set of training examples. Each training example includes a training network input and, for each training network input, a ground truth output sequence, i.e., the output sequence that the sequence generation neural network 110 should generate by processing the training network input. For example, for speech recognition, each training network input represents a utterance and the ground truth output sequence for a given training network input is the transcription of the utterance represented by the given training network input. As another example, for machine translation, each training network input is text in a source language and the ground truth output sequence for a given training network input is the translation of the text in the source language into a target language.

[0030] At each iteration of the training process, the training engine 120 in the system 100 applies a parameter value update 116 to the current network parameter values 114 since the iteration.

[0031] In particular, at each iteration, the training engine 120 or more generally the system 100 causes the sequence generation neural network 110 to generate a batch 142 of new output sequences in accordance with the current network parameter values 114, i.e., by using the sequence generation neural network 110 and processing each training network input in a batch of training network inputs 132 in accordance with the current parameter values 114 to map the network input to a new output sequence.

[0032] Each new output sequence in the batch includes a respective system output from the vocabulary of system outputs at each of a plurality of output positions. As described above, the neural network 110 generates a likelihood distribution over the vocabulary of system outputs at each of the plurality of output positions and then selects, e.g., samples, a system output from the vocabulary in accordance with the likelihood distribution to generate the output sequence.

[0033] The Q-value engine 140 then uses the ground truth output sequences 162 of the training network inputs in the batch to determine Q-values 144 for each new output sequence in the batch. In particular, the Q-value engine 140 generates respective Q-values for each possible system output in the vocabulary for each position in a given output sequence.

[0034] The Q-value of a particular possible system output at a given position in a given output sequence is the highest possible quality score that can be assigned to any candidate output sequence that: (i) starts with the prefix of the system output at positions before the given output position included in the given output sequence and (ii) has the particular possible system output immediately following that prefix. That is, the candidate output sequence can have any possible suffix, as long as the suffix is immediately preceded by (i) the prefix and (ii) the particular possible system output. In other words, the candidate output sequence can be any sequence of the form [p,a,s], where p is the prefix, a is the particular possible system output, and s is any suffix of zero or more possible system outputs. The quality score of the candidate output sequence measures the quality of the candidate output sequence relative to the corresponding ground truth output sequence - i.e., the ground truth output sequence of the network input from which the given output sequence was generated.

[0035] The generation of the Q-value is described in more detail below with reference to Figures 2 - 5 generate the Q-value.

[0036] The training engine 120 uses the Q-value 144 and the current likelihood distribution 152 generated by the neural network as part of the batch 142 of new output sequences to determine the parameter update 116, which is then applied, e.g., added, to the current value 114 of the network parameters to generate an updated network parameter value. The use of the Q-value to determine the parameter update is described below with reference to Figure 2 determine the parameter update.

[0037] By iteratively updating the network parameters in this way, the system 100 can effectively train the sequence generation neural network 110 to generate high-quality output sequences.

[0038] Although Figure 1 only a single training engine 120 communicating with a single instance of the sequence generation neural network 110 is shown, in some implementations the training process can be distributed across multiple hardware devices. In particular, to accelerate training, an asynchronous or synchronous distributed setup can be employed, where a parameter server stores the shared model parameters for many copies of the sequence generation neural network. The training engine for each network copy asynchronously or synchronously samples a batch of sequences from its local network copy and computes the following gradients. The gradients are then sent to the parameter server, which updates the shared parameters. The copies periodically update their local parameter values with the latest parameters from the parameter server.

[0039] Figure 2It is a flowchart of an example process 200 for training a sequence generation neural network system. For convenience, process 200 is described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed neural network training system, such as Figure 1 the neural network training system 100, is capable of performing process 200.

[0040] The system is capable of performing process 200 for each training example in a batch of training examples to determine a corresponding parameter update for each training example in the batch. The batch typically includes a fixed number of training examples, such as ten, fifty, or one hundred. The system can then generate a final parameter update for the batch, for example, by averaging or summing the parameter updates of the training examples, and then apply the final parameter update, for example, add it, to the current value of the parameters to generate an updated parameter value.

[0041] The system uses a sequence generation neural network to process the training network input in the training examples according to the current value of the network parameters to generate a new output sequence, that is, to map the training network input to a new output sequence (step 202). To generate a new output sequence, the system samples according to the likelihood distribution generated by the sequence generation neural network, for example, until a pre-determined sequence end output token is sampled or until the sequence reaches a pre-determined maximum length.

[0042] The system generates a Q value for each of the possible system outputs in the vocabulary for each position in the new output sequence (step 204). Generating the Q value for a given position in the output sequence is described below with reference to Figures 3 - 5 description.

[0043] The system determines an update to the current value of the network parameters for each of the positions, which increases the likelihood that the neural network generates a system output with a high quality score at that position (step 206). That is, the system generates an update that makes it more likely for the neural network to sample a system output with a high quality score at that position.

[0044] To determine the update for a given position, the system transforms the Q values of the possible system outputs at the given position into a target likelihood distribution over the possible system outputs in the vocabulary.

[0045] For example, the system can generate a target likelihood distribution by applying softmax to the Q values of the possible system outputs to generate a corresponding likelihood for each of the possible system outputs. In some implementations, softmax is applied at a reduced temperature.

[0046] In other words, the system can generate a likelihood for a possible system output a in the vocabulary by applying the following transformation:

[0047]

[0048] where is the Q-value of token a, the sum is over all tokens a' in the vocabulary, and τ is the temperature parameter. To apply softmax at a reduced temperature, the system sets the temperature parameter to a value between zero and one. In particular, in some implementations, the system sets the temperature parameter to a value approaching the limit τ → 0, i.e., a very small positive value, to result in a "hard" distribution with one or more very strong peaks, i.e., a distribution where all probabilities except for the probabilities of a small subset of outputs in the vocabulary are approximately zero.

[0049] The system then determines an update for a given position by computing the gradient of the network parameters with respect to the objective function and then determining an update to the parameters based on that gradient, where the objective function depends on the divergence between the target likelihood distribution at the output position and the likelihood distribution generated by the neural network for the output position.

[0050] For example, the objective function can be the Kullback-Leibler (KL) divergence between the target likelihood distribution at the output position and the likelihood distribution generated by the neural network for the output position.

[0051] The system is able to determine an update to the parameters based on the gradient by applying an update rule that defines how to map the gradient to an update of the parameter values, such as the rmsProp update rule, the Adam update rule, or the stochastic gradient descent update rule.

[0052] The system determines an update to the current value of the training example from the updates determined for each of the multiple positions (step 208). For example, the system is able to sum or average the updates at each position to determine an update to the current value of the training example.

[0053] Figure 3 is a flowchart of an example process 300 for determining the Q-value of a given output position in an output sequence. For convenience, process 300 is described as being performed by a system of one or more computers located at one or more positions. For example, a neural network training system, such as Figure 1 neural network training system 100, suitably programmed, is capable of performing process 100.

[0054] The system is able to perform process 300 for each of the output positions in the output sequence generated during the training of the sequence generation neural network.

[0055] The system identifies the prefix of the system output at a position before a particular output position in the output sequence (step 302). In other words, the system identifies as the prefix at that position a partial output sequence that consists of the system outputs in the output sequence at positions before a given position in the output sequence. For the first position in the output sequence, the prefix is the empty set, i.e., there is no output in the prefix.

[0056] The system generates a corresponding Q value for each possible system output in the vocabulary (step 304).

[0057] In particular, the system determines the highest quality score that can be assigned to any one of the candidate output sequences from a group of possible candidate output sequences, where the possible candidate output sequences include: (i) the identified prefix, which is followed by (ii) a possible system output and is followed by (iii) a suffix that is any one of zero or more system outputs. That is, the group of possible candidate output sequences all start with the same identified prefix followed by the same system output, but all have different suffixes. The system sets the Q value to the determined highest quality score.

[0058] The quality score of a given candidate output sequence measures the quality of the given candidate output sequence relative to the ground truth output sequence. That is, the quality score measures the difference between the candidate output sequence and the ground truth output sequence according to a quality metric. Generally, the metric used to evaluate this quality depends on the type of sequence generated by the neural network.

[0059] As a particular example, when the output sequence is a natural language sequence and the possible outputs in the vocabulary are sequences of natural language characters (optionally augmented with one or more special characters, such as a whitespace symbol representing the space between characters and a sequence end symbol representing that the output sequence should terminate), the metric can be based on the edit distance between the candidate output sequence and the ground truth output sequence.

[0060] The edit distance between two sequences u and v is the minimum number of insert, delete, and substitute edits required to transform u into v and vice versa. Thus, when the quality metric is based on the edit distance, the highest quality score that can be assigned is the quality score of the candidate output sequence that has the minimum edit distance from the ground truth output sequence.

[0061] As a particular example, the quality metric can be the negative of the edit distance or can be proportional to one (or another positive constant) plus the reciprocal of the edit distance.

[0062] In the following, reference is made to Figure 4 an example showing the identification of the edit distance.

[0063] In the following, reference is made to Figure 5Describe techniques for efficiently identifying the highest quality scores when the distance metric is based on edit distance.

[0064] Figure 4 Show an example of applying the OCD training technique when the quality metric is based on edit distance.

[0065] In particular, Figure 4 show a ground truth output sequence (referred to as the target sequence) and a new output sequence generated by a neural network (referred to as the generated sequence). In Figure 4 the example, the ground truth output sequence is "as_he_talks_his_wife", while the new output sequence is "as_ee_talks_whose_wife".

[0066] For each position in the new output sequence, Figure 4 also show the optimal extension for the edit distance, i.e., the possible system output that will have the highest Q value for that position and thus the highest probability in the target likelihood distribution generated for that position.

[0067] In Figure 4 the example, the optimal extension for a given output position is shown below and immediately to the left of the output at the given output position in the new output sequence.

[0068] As a specific example, for the first position in the output sequence (with output "a"), the optimal extension is "a" because the prefix for the first position is empty and an edit distance of zero can be achieved by matching the first output ("a") in the ground truth output sequence.

[0069] As another specific example, for the fifth position in the output sequence (where the output sequence is "as_ee" and the prefix would be "as_e"), there are three optimal extensions "e", "h", and "_". This is because any of these three system outputs following the prefix "as_e" (when combined with the appropriate suffix) can produce a candidate output sequence with an edit distance of one. Thus, each of these three possible system outputs will receive the same Q value, and the target likelihood distribution for the fifth position will assign the same likelihood to each of these three possible system outputs.

[0070] Figure 5 is a flowchart of an example process 500 for determining the highest quality score for a specific system output with a specific prefix when the quality metric is based on edit distance. For convenience, process 500 is described as being performed by a system of one or more computers located at one or more positions. For example, a neural network training system that is appropriately programmed, such as Figure 1The neural network training system 100 is capable of performing process 500.

[0071] Specifically, the system is capable of performing process 500 to efficiently determine the highest quality score that can be assigned to any candidate output sequence that includes a specific word head followed by a specific possible system output and followed by the end of any one of one or more system outputs.

[0072] The system determines the highest quality score that can be assigned to any candidate output sequence that includes a specific word head (from step 302) followed by any ground truth end that is part of the ground truth output sequence (step 502). In other words, given word head p, the system determines the highest quality score that can be assigned to any candidate output sequence [p, s], where any candidate output sequence [p, s] is the concatenation of a specific word head p and any end s that is part of the ground truth output sequence.

[0073] The system identifies one or more ground truth word heads of the ground truth output sequence, and the specific word head has the highest quality score relative to the one or more ground truth word heads of the ground truth output sequence (step 504). In other words, the system identifies one or more ground truth word heads of the ground truth output sequence, and the specific word head has the smallest edit distance relative to the one or more ground truth word heads of the ground truth output sequence.

[0074] For each of the identified ground truth word heads, the system identifies the corresponding ground truth end that follows the identified ground truth word head in the ground truth sequence (step 506).

[0075] The system determines whether the specific possible system output is the first system output in any of the identified ground truth ends (step 508).

[0076] If the system output is the first system output in one or more of the identified ground truth ends, the system assigns the highest quality score that can be assigned to any candidate output sequence that includes the word head, i.e., the highest quality score determined in step 502, as the highest quality score of the specific possible system output, where the word head is followed by any ground truth end that is part of the ground truth output sequence (step 510).

[0077] If the system output is not the first system output in any of the identified ground truth ends, the system determines the highest quality score that can be assigned to any candidate output sequence that includes a specific word head followed by a possible system output that is not the first system output in any of the identified ends and followed by any ground truth end that is part of the ground truth output sequence (step 512).

[0078] The system assigns as the highest quality score for a particular possible system output the highest quality score that can be assigned to any candidate output sequence that includes a word head that is followed by a possible system output that is not one of the recognized ground truth word tails and that is followed by any ground truth word tail that is part of the ground truth output sequence (step 514).

[0079] By using procedure 500 to identify the highest quality score for possible system outputs, the system is able to compute the highest quality score for each word head and for each possible system output using dynamic programming at a complexity of O(|y’|*|y|+|V|*|y|), where |y’| is the number of outputs in the ground truth sequence, |y| is the number of outputs in the generated output sequence, and |V| is the number of outputs in the vocabulary. Thus, the system can perform this search for quality scores without impeding the training process, i.e., without significantly affecting the running time of a given training iteration.

[0080] Procedure 500 is depicted as the pseudocode of the dynamic programming algorithm in Table 2 below. In particular, the pseudocode in Table 2 refers to the ground truth sequence as the reference sequence r and the new output sequence as the hypothesis sequence h.

[0081]

[0082] Table 2

[0083] This specification, together with the system and computer program components, uses the term "configured". For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed thereon software, firmware, hardware, or a combination of software, firmware, and hardware that in operation causes the system to perform those operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operations or actions.

[0084] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of them. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus.

[0085] The term “data processing apparatus” refers to data processing hardware and includes all kinds of devices, equipment, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus may also be or further include special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus may optionally include, in addition to hardware, code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0086] A computer program, which may also be referred to as or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language including a compiled or interpreted language, or a declarative or procedural language; and it can be deployed in any form, including as a stand-alone program or as a module, a component, a subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a part of a file that holds other programs or data, e.g., one or more scripts in a markup language document; in a single file dedicated to the program; or in multiple coordinated files, e.g., files that store one or more modules, subroutines, or portions of code. A computer program may be deployed to execute on one computer or on multiple computers interconnected by a data communication network and located at one site or distributed across multiple sites.

[0087] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or at all, and it can be stored on a storage device in one or more locations. Thus, for example, an indexing database can include multiple collections of data, each of which can be organized and accessed differently.

[0088] Similarly, in this specification the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. In general, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same computer or multiple computers.

[0089] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, a special-purpose logic circuit such as an FPGA or ASIC, or by a combination of special-purpose logic circuits and one or more programmed computers.

[0090] Computers suitable for executing computer programs can be based on a general or special purpose microprocessor or both, or any other kind of central processing unit. In general, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a central processing unit for executing or carrying out instructions and one or more storage devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuits. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to the one or more mass storage devices, or both, for storing data. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game controller, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, etc.

[0091] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including by way of example semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0092] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device by which the user can provide input to the computer, the display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and the pointing device such as a mouse or a trackball. Other kinds of devices can also be used to provide for interaction with a user; for example, feedback provided to the user can be any form of sensory feedback such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from devices used by the user; for example, by sending a web page to a web browser on a user's device in response to a request received from the web browser. Further, a computer can interact with a user by sending text messages or other forms of messages to a personal device and then receiving a response message from the user, the personal device such as a smart phone running a messaging application.

[0093] The data processing apparatus for implementing a machine learning model may further include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive portions of machine learning training or production - i.e., inference, workloads.

[0094] A machine learning framework can be used to implement and deploy a machine learning model, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.

[0095] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component such as a data server; or a middleware component such as an application server; or a frontend component such as a client computer having a graphical user interface, a web browser, or an app by which a user can interact with an implementation of the subject matter described in this specification; or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication such as, for example, a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN) such as the Internet.

[0096] A computing system may include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship of the client and the server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, the server transmits data such as HTML pages to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as the client and receiving user input from the user. Data generated at the user device may be received at the server, for example, the result of a user interaction.

[0097] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0098] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.

[0099] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order and still achieve the desired result. As one example, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method for training a neural network, the neural network having a plurality of network parameters and being configured to map a system input to an output sequence including a plurality of system outputs, wherein, The method includes: obtaining a batch of training examples, each training example including a training network input and, for each training network input, a ground truth output sequence; for each training example in the batch of training examples: processing the training network input in the training example using the neural network according to a current value of the network parameters to map the training network input to an output sequence, the output sequence including a respective system output from a vocabulary of possible system outputs at each of a plurality of output positions, wherein the possible system outputs in the vocabulary include tokens of natural language, and wherein the neural network is configured to, for each output position in the plurality of output positions, generate a likelihood distribution over the possible system outputs in the vocabulary and use the likelihood distribution to select the system output at that output position; for each particular output position in the plurality of output positions in the output sequence: identifying a head word of a system output at a position before the particular output position included in the output sequence, for each possible system output in the vocabulary, determining a highest quality score that can be assigned to any candidate output sequence including the head word followed by the possible system output and followed by any tail word of one or more system outputs, wherein the quality score measures the quality of the candidate output sequence relative to the ground truth output sequence, and using the highest quality score of the possible system output to determine an update to the current value of the network parameters, the update increasing the likelihood that the neural network generates a system output with a high quality score at that position, including: generating a target likelihood distribution for the output position based on the highest quality score of the possible system output, determining a gradient of the network parameters with respect to an objective function that depends on a divergence between the target likelihood distribution for the output position and the likelihood distribution generated by the neural network for the output position, and using the gradient to determine the update to the current value; and determining an updated value of the network parameters based on the updates for the particular output position in the output sequence generated by the neural network for the batch of training examples.

2. The method according to claim 1, further comprising: outputting the trained neural network for use in mapping a new network input to a new output sequence.

3. The method according to claim 1, wherein, Generating the target likelihood distribution includes applying softmax to the highest quality score of the possible system output to generate a respective likelihood for each of the possible system outputs.

4. The method according to claim 3, wherein, The softmax is applied at a reduced temperature.

5. The method according to claim 4, wherein The quality score is based on an edit distance between the candidate output sequence and the ground truth output sequence, and wherein the highest quality score that can be assigned is the quality score of the candidate output sequence having the smallest edit distance from the ground truth output sequence.

6. The method according to any one of claims 1-5, wherein, Determining the highest quality score assignable to any candidate output sequence that includes the headword, the headword being followed by the possible system output and by any tail of one or more system outputs, includes: Determining the highest quality score assignable to any candidate output sequence that includes the headword, the headword being followed by any ground truth tail that is part of the ground truth output sequence; Identifying one or more ground truth headwords of the ground truth output sequence, the headwords having the highest quality score relative to one or more ground truth headwords of the ground truth output sequence, For each of the identified ground truth headwords, identifying the corresponding ground truth tail in the ground truth sequence that follows the identified ground truth headword, and When the possible system output is the first system output in any of the identified ground truth tails, assigning the highest quality score assignable to any candidate output sequence that includes the headword, the headword being followed by any ground truth tail that is part of the ground truth output sequence, as the highest quality score of the possible system output.

7. The method according to claim 6, wherein, Determining the highest quality score assignable to any candidate output sequence that includes the headword, the headword being followed by the possible system output and by any tail of one or more system outputs, includes: Determining the highest quality score assignable to any candidate output sequence that includes the headword, the headword being followed by a possible system output that is not the first system output in any of the identified tails and by any ground truth tail that is part of the ground truth output sequence; When the possible system output is not the first system output in any of the identified ground truth tails, assigning the highest quality score assignable to any candidate output sequence that includes the headword, the headword being followed by a possible system output that is not the first system output in any of the identified ground truth tails and by any ground truth tail that is part of the ground truth output sequence, as the highest quality score of the possible system output.

8. The method according to claim 1, wherein The neural network is one or more of a machine translation neural network, a speech recognition neural network, and a natural language processing neural network.

9. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 - 8.

10. A computer storage medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 - 8.