Attention-based sequence transduction neural network
By employing an attention-based encoder-decoder neural network system, the system addresses the inefficiencies of recurrent neural networks in sequence conversion, achieving faster training and inference with improved accuracy in sequence transformation tasks.
Patent Information
- Application Number
- JP2025018368
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-08-04
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2038-05-23
AI Technical Summary
Existing sequence conversion methods using recurrent neural networks face challenges with long training and inference times due to their sequential nature, leading to high computational resource utilization.
The implementation of an attention-based encoder-decoder neural network system that transforms input sequences into output sequences, eliminating the need for recurrent neural network layers and enabling easier parallelization.
This approach reduces the number of operations required to associate signals from different positions in the sequence to a constant, allowing for faster training and inference, and improves the accuracy of sequence transformation tasks such as machine translation.
Smart Images

Figure 2025084774000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a regular application of U.S. Provisional Patent Application No. 62 / 510,256, filed on May 23, 2017, and U.S. Provisional Patent Application No. 62 / 541,594, filed on August 4, 2017, and claims the priority thereof. The entire contents of the foregoing applications are incorporated herein by reference.
[0002] This specification relates to using neural networks to transform sequences.
Background Art
[0003] A neural network is a machine - learning model that employs one or more layers of non - linear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as an input to the next layer within the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of each set of parameters.
Summary of the Invention
Means for Solving the Problems
[0004] This specification describes a system implemented as a computer program on one or more computers that generates an output sequence including respective outputs at a plurality of positions in output order from an input sequence including respective inputs at a plurality of positions in input order, i.e., transforms the input sequence into the output sequence. In particular, the system uses an encoder neural network and a decoder neural network, both of which are attention - based, to generate the output sequence.
[0005] Certain embodiments of the subject matter described herein may be practiced to realize one or more of the following advantages.
[0006] Many existing approaches to sequence conversion using neural networks use recurrent neural networks in the encoder and decoder. These types of networks can achieve good performance in sequence conversion tasks, but their computations are inherently sequential, i.e., a recurrent neural network generates an output at the current time step conditioned on the hidden state of the recurrent neural network at the previous time step. This sequential nature hinders parallelization and as a result leads to long training and inference times and a correspondingly large computational resource utilization workload.
[0007] On the other hand, since the encoder and decoder of the described sequence conversion neural network are attention-based, the sequence conversion neural network can convert sequences more quickly and can be trained more rapidly, or both because the operation of the network can be more easily parallelized. That is, since the described sequence conversion neural network relies entirely on the attention mechanism to elicit the global dependency relationship between the input and output and does not employ any recurrent neural network layers, the problems associated with the long training and inference times and high resource usage caused by the sequential nature of the recurrent neural network layers are alleviated.
[0008] Furthermore, even if the training and inference times are short, the sequence transformation neural network can transform sequences more accurately than existing networks based on convolutional or recurrent layers. In particular, in conventional models, the number of operations required to associate signals from two arbitrary input or output positions increases with the distance between positions, which depends linearly or logarithmically on the model architecture, for example. This makes it even more difficult to learn dependencies between remote positions during training. In the sequence transformation neural network currently being described, this number of operations is reduced to a constant number of operations through the use of attention (and, in particular, self-attention), without relying on recursion or convolution. Self-attention, sometimes referred to as intra-attention, is an attention mechanism that relates different positions of a single sequence in order to compute a representation of the sequence. By using the attention mechanism, the sequence transformation neural network can effectively learn dependencies between remote positions during training, and can improve the accuracy of the sequence transformation neural network in various transformation tasks such as machine translation. In fact, the described sequence transformation neural network is easier to train than conventional machine translation neural networks and can achieve state-of-the-art results in the machine translation task despite generating outputs quickly. The sequence transformation neural network can also exhibit improved performance over conventional machine translation neural networks without performing task-specific adjustments through the use of the attention mechanism.
[0009] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010]
Figure 1
Figure 2
Figure 3
[0011] In various drawings, like numbers and symbols denote like elements.
[0012] This specification describes a system implemented as a computer program on one or more computers in one or more locations that generates an output sequence including respective outputs at each of a plurality of positions in output order from an input sequence including respective inputs at each of a plurality of positions in input order, that is, converts the input sequence into the output sequence.
[0013] For example, the system may be a neural machine translation system. That is, when the input sequence is a sequence of words in a source language, such as a sentence or a phrase, the output sequence may be a conversion of the input sequence into a target language, that is, a sequence of words in the target language representing the sequence of words in the source language.
[0014] As another example, the system may be an automatic speech recognition system. That is, when the input sequence is a sequence of audio data representing an oral utterance, the output sequence may be a sequence of graphemes, features, or words representing the utterance, that is, the orthographic transcription of the input sequence.
[0015] As another example, the system may be a natural language processing system. For example, if the input sequence is a sequence of words in a source language, such as a sentence or a phrase, the output sequence may be a summary of the input sequence in the source language, that is, a sequence having fewer words than the input sequence but retaining the essential meaning of the input sequence. As another example, if the input sequence is a sequence of words forming a question, the output sequence may be a sequence of words forming an answer to the question.
[0016] As another example, the system may be part of a computer-aided medical diagnosis system. For example, the input sequence may be a sequence of data from an electronic medical record, and the output sequence may be a sequence of predicted treatments.
[0017] As another example, the system may be part of an image processing system. For example, the input sequence may be an image, that is, a sequence of brightness values from an image, and the output may be a sequence of text describing the image. As another example, the input sequence may be a sequence of text or different contexts, and the output sequence may be an image explaining the context.
[0018] In particular, the neural network includes an encoder neural network and a decoder neural network. Generally, both the encoder and the decoder are attention-based, that is, both apply an attention mechanism across their received inputs while converting the input sequence. In some cases, neither the encoder nor the decoder includes a convolutional layer or a recurrent layer.
[0019] FIG. 1 shows an exemplary neural network system 100. The neural network system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations where the systems, components, and techniques described below may be implemented.
[0020] The neural network system 100 receives an input sequence 102 and processes the input sequence 102 to convert the input sequence 102 into an output sequence 152.
[0021] The input sequence 102 has respective network inputs at each of a plurality of input positions in input order, and the output sequence 152 has respective network outputs at each of a plurality of output positions in output order. That is, the input sequence 102 has a plurality of inputs arranged in input order, and the output sequence 152 has a plurality of outputs arranged in output order.
[0022] As described above, the neural network system 100 can perform any of a variety of tasks that require processing sequential inputs to generate sequential outputs.
[0023] The neural network system 100 includes an attention-based sequence conversion neural network 108, and this neural network 108 includes an encoder neural network 110 and a decoder neural network 150.
[0024] The encoder neural network 110 is configured to receive the input sequence 102 and generate an encoded representation of each respective network input within the input sequence. Generally, the encoded representation is a vector of numerical values or other ordered collection.
[0025] Next, decoder neural network 150 is configured to use the encoded representation of the network input to generate output sequence 152.
[0026] Generally, and as described in more detail below, both encoder 110 and decoder 150 are attention-based. In some cases, neither the encoder nor the decoder includes convolutional layers or recurrent layers.
[0027] Encoder neural network 110 includes embedding layer 120 and a sequence of one or more encoder sub-networks 130. In particular, as shown in FIG. 1, the encoder neural network includes N encoder sub-networks 130.
[0028] Embedding layer 120 is configured to map each network input in the input sequence to a numerical representation of the network input in an embedding space, for example, a vector in the embedding space. Embedding layer 120 then provides the numerical representation of the network input to the first sub-network in the sequence of encoder sub-networks 130, that is, the first encoder sub-network 130 of the N encoder sub-networks 130.
[0029] In particular, in some embodiments, the embedding layer 120 maps each network input to an embedded representation of the network input and then combines the embedded representation of the network input with the positional embedding of the input position of the network input in the order of input, for example, by summing or averaging, to generate a combined embedded representation of the network input. That is, each position in the input sequence has a corresponding embedding, and for each network input, the embedding layer 120 combines the embedded representation of the network input with the embedding of the position of the network input in the input sequence. Such positional embeddings can enable the model to fully utilize the order of the input sequence without relying on recurrence or convolution.
[0030] In some cases, the positional embeddings are learned. As used herein, the term "learned" means that the operation or value is being adjusted during the training of the sequence transformation neural network 108. The training of the sequence transformation neural network 108 is described below with reference to FIG. 3.
[0031] In some cases, the positional embeddings are fixed and different for each position. For example, the embeddings may be composed of sine and cosine functions of different frequencies and can satisfy the following equation.
[0032]
Equation
[0033] where pos is the position, i is the dimension within the positional embedding, and d model is the number of dimensions of the positional embedding (and other vectors processed by the neural network 108). The use of sine function positional embeddings enables the model to extrapolate to longer sequence lengths, thereby increasing the range of applications in which the model can be employed.
[0034] Next, the combined embedded representation is used as the numerical representation of the network input.
[0035] Each of the encoder sub-networks 130 is configured to receive respective encoder sub-network inputs for each of a plurality of input positions and to generate respective sub-network outputs for each of the plurality of input positions.
[0036] Next, the encoder sub-network output generated by the last encoder sub-network in the sequence is used as the encoded representation of the network input.
[0037] For the first encoder sub-network in the sequence, the encoder sub-network input is the numerical representation generated by the embedding layer 120, and for each encoder sub-network other than the first encoder sub-network in the sequence, the encoder sub-network input is the encoder sub-network output of the preceding encoder sub-network in the sequence.
[0038] Each encoder sub-network 130 includes an encoder self-attention sub-layer 132. The encoder self-attention sub-layer 132 receives sub-network inputs for each of a plurality of input positions and, for each particular input position in input order, applies an attention mechanism across the encoder sub-network inputs using one or more queries derived from the encoder sub-network inputs at the particular input position to generate respective outputs for each of the particular input positions. In some cases, the attention mechanism is a multi-head attention mechanism. The attention mechanism and how the attention mechanism is applied by the encoder self-attention sub-layer 132 are described in more detail below with reference to FIG. 2.
[0039] In some embodiments, each of the encoder sub-networks 130 also includes a residual connection layer that combines the output of the encoder self-attention sub-layer with the input to the encoder self-attention sub-layer to produce an encoder self-attention residual output, and a layer normalization layer that applies layer normalization to the encoder self-attention residual output. These two layers are collectively referred to as the "Add & Norm" operation in Figure 1.
[0040] Some or all of the encoder sub-network can also include a per-position feed-forward layer 134 configured to operate at each position within the input sequence. In particular, for each input sequence position, the feed-forward layer 134 is configured to receive the input at the input position and apply a sequence of transformations to the input at the input position to produce an output at the input position. For example, the sequence of transformations can include two or more learned linear transformations, each split by an activation function, such as an activation function per non-linear element, such as the ReLU activation function, which can enable faster and more effective training on large and complex databases. The input received by the per-position feed-forward layer 134 can be the output of the layer normalization layer if residual and layer normalization layers are included, or the output of the encoder self-attention sub-layer 132 if residual and layer normalization layers are not included. The transformations applied by layer 134 are generally the same for each input position (however, different feed-forward layers of different sub-networks apply different transformations).
[0041] When the encoder sub-network 130 includes a feed-forward layer 134 for each position, the encoder sub-network may also include a residual connection layer that combines the output of the feed-forward layer for each position with the input to the feed-forward layer for each position to generate a residual output for each encoder position, and a layer normalization layer that applies layer normalization to the residual output for each encoder position. These two layers are also collectively referred to as the "Add and Normalize" operation in FIG. 1. Then, the output of this layer normalization layer may be used as the output of the encoder sub-network 130.
[0042] When the encoder neural network 110 generates an encoded representation, the decoder neural network 150 is configured to generate an output sequence in an autoregressive manner.
[0043] That is, in each of a plurality of generation time steps, the decoder neural network 150 generates the network output of the corresponding output position conditioned on (i) the encoded representation and (ii) the network output at the output positions preceding the output position in the output order, thereby generating an output sequence.
[0044] In particular, for a given output position, the decoder neural network generates an output that defines a probability distribution over the possible network outputs at the given output position. Then, the decoder neural network can select the network output of the output position by sampling from the probability distribution or by selecting the network output with the highest probability.
[0045] Since the decoder neural network 150 is autoregressive, at each generation time step, the decoder 150 operates on the network output that has already been generated prior to the generation time step, i.e., the network output at the output positions preceding the output position corresponding to the order of output. In some embodiments, to ensure that this holds during inference and training, at each generation time step, the decoder neural network 150 shifts the already generated network output by only one output order position to the right (i.e., introduces one position offset into the already generated network output sequence) and masks certain operations so that it can pay attention only to positions up to and including that position (not subsequent positions) within the output sequence, as will be described in more detail below. The remainder of the following description explains that when generating a given output at a given output position, the various components of the decoder 150 operate on the data at the output positions preceding the given output position (and not on the data at any other output position), but it will be understood that this type of conditioning can be effectively implemented using the shift described above.
[0046] The decoder neural network 150 includes an embedding layer 160, a sequence of decoder sub-networks 170, a linear layer 180, and a softmax layer 190. In particular, as shown in FIG. 1, the decoder neural network includes N decoder sub-networks 170. However, the example of FIG. 1 shows an encoder 110 and a decoder 150 that include the same number of sub-networks, but in some cases, the encoder 110 and the decoder 150 may include different numbers of sub-networks. That is, the decoder 150 can include more or fewer sub-networks than the encoder 110.
[0047] The embedding layer 160 is configured to map, for each network output at an output position preceding the current output position in output order at each generation time step, the network output to a numerical representation of the network output within the embedding space. The embedding layer 160 then provides the numerical representation of the network output to the first sub-network 170 within the sequence of decoder sub-networks, i.e., the first decoder sub-network 170 of the N decoder sub-networks.
[0048] In particular, in some embodiments, the embedding layer 160 is configured to map each network output to an embedded representation of the network output, combine the embedded representation of the network output with the positional embedding of the output position of the network output in output order, and generate a combined embedded representation of the network output. The combined embedded representation is then used as the numerical representation of the network output. The embedding layer 160 generates the combined embedded representation in the same way as described above with reference to the embedding layer 120.
[0049] Each decoder sub-network 170 is configured to receive, at each generation time step, a respective decoder sub-network input for each of a plurality of output positions preceding the corresponding output position, and generate a respective decoder sub-network output for each of a plurality of output positions preceding the corresponding output position (or equivalently, when the output sequence is shifted to the right, the network outputs at and including the current output position).
[0050] In particular, each decoder sub-network 170 includes two different attention sub-layers: a decoder self-attention sub-layer 172 and an encoder-decoder attention sub-layer 174.
[0051] At each generation time step, each decoder self-attention sub-layer 172 receives inputs for each output position preceding the corresponding output position, and for each of the specific output positions, uses one or more queries derived from the inputs at the specific output position to apply an attention mechanism across the inputs at the output positions preceding the corresponding position to generate an updated representation of the specific output position. That is, the decoder self-attention sub-layer 172 applies an attention mechanism that is masked so as not to attend to or process any data that is not at positions preceding the current output position in the output sequence.
[0052] On the other hand, at each generation time step, each encoder-decoder attention sub-layer 174 receives inputs for each output position preceding the corresponding output position, and for each of the output positions, uses one or more queries derived from the inputs at the output position to apply an attention mechanism across the representations encoded at the input positions to generate an updated representation of the output position. Thus, the encoder-decoder attention sub-layer 174 applies attention across the encoded representations, while the decoder self-attention sub-layer 172 applies attention across the inputs at the output position.
[0053] The attention mechanism applied by each of these attention sub-layers will be described in more detail below with reference to Figure 2.
[0054] In Figure 1, the decoder self-attention sub-layer 172 is shown as being before the encoder-decoder attention sub-layer in the processing order within the decoder sub-network 170. However, in other examples, the decoder self-attention sub-layer 172 may be after the encoder-decoder attention sub-layer 174 in the processing order within the decoder sub-network 170, or different sub-networks may have different processing orders.
[0055] In some embodiments, each decoder subnetwork 170 includes a residual connection layer that combines the output of an attention sublayer with the input to the attention sublayer to produce a residual output after the decoder self-attention sublayer 172, after the encoder-decoder attention sublayer 174, or after each of the two sublayers, and a layer normalization layer that applies layer normalization to the residual output. FIG. 1 shows these two layers, both referred to as "add and normalize" operations, inserted after each of the two sublayers.
[0056] Some or all of the decoder subnetwork 170 also includes a per-position feed-forward layer 176 configured to operate in a manner similar to the per-position feed-forward layer 134 from the encoder 110. In particular, layer 176 is configured to receive an input at an output position for each output position preceding the corresponding output position at each generation time step and apply a sequence of transformations to the input at the output position to produce an output at the output position. For example, the sequence of transformations can include two or more learned linear transformations each divided by an activation function, such as an activation function per non-linear element, such as a ReLU activation function. The input received by the per-position feed-forward layer 176 can be the output of the layer normalization layer (following the last attention sublayer within the subnetwork 170) if residual and layer normalization layers are included, or the output of the last attention sublayer within the subnetwork 170 if residual and layer normalization layers are not included.
[0057] In the case where the decoder sub-network 170 includes a feed-forward layer 176 for each position, the decoder sub-network may also include a residual connection layer that combines the output of the feed-forward layer for each position with the input to the feed-forward layer for each position to generate a residual output for each decoder position, and a layer normalization layer that applies layer normalization to the residual output for each decoder position. These two layers are also collectively referred to as the "addition and normalization" operation in FIG. 1. Then, the output of this layer normalization layer may be used as the output of the decoder sub-network 170.
[0058] At each generation time step, the linear layer 180 applies a learned linear transformation to the output of the last decoder sub-network 170 in order to project the output of the last decoder sub-network 170 into a space appropriate for processing by the softmax layer 190. The softmax layer 190 then applies the softmax function across the output of the linear layer 180 to generate a probability distribution across the possible network outputs at the generation time step. As described above, the decoder 150 can then select a network output from the possible network outputs using the probability distribution.
[0059] FIG. 2 is a diagram 200 showing the attention mechanism applied by the attention sub-layer in the sub-networks of the encoder neural network 110 and the decoder neural network 150.
[0060] Generally, the attention mechanism maps a set of queries and key-value pairs to an output, where the query, key, and value are all vectors. The output is calculated as a weighted sum of the values, where the weight assigned to each value is calculated by a compatibility function of the query with the corresponding key.
[0061] More specifically, each attention sublayer applies a scaled dot - product attention mechanism 230. In scaled dot - product attention, for a given query, the attention sublayer calculates the dot - product of the query with all of the keys, divides each of the dot - products by a scaling factor, e.g., by the square root of the dimensionality of the query and the keys, and applies the softmax function over the scaled dot - products to obtain weights for the values. The attention sublayer then calculates the weighted sum of the values according to these weights. Thus, in the case of scaled dot - product attention, the scoring function is the dot - product, and the output of the scoring function is further scaled by the scaling factor.
[0062] During operation, and as shown on the left side of FIG. 2, the attention sublayer calculates attention simultaneously over a set of queries. In particular, the attention sublayer packs the queries into a matrix Q, the keys into a matrix K, and the values into a matrix V. To pack a set of vectors into a matrix, the attention sublayer can generate a matrix that contains the vectors as the rows of the matrix.
[0063] The attention sublayer then performs a matrix multiplication (MatMul) between matrix Q and the transpose of matrix K to generate a matrix of scoring function outputs.
[0064] The attention sublayer then scales the scoring function output matrix, i.e., divides each element of the matrix by the scaling factor.
[0065] The attention sublayer then applies softmax to the scaled output matrix to generate a matrix of weights, and performs a matrix multiplication (MatMul) between the weight matrix and matrix V to generate an output matrix that contains the output of the attention mechanism for each value.
[0066] In the case of the sublayer that uses masking, i.e., the decoder attention sublayer, the attention sublayer masks the reduced output matrix before applying the softmax. That is, the attention sublayer excludes (sets to negative infinity) all values of the reduced output matrix corresponding to positions after the current output position by means of the mask.
[0067] In some embodiments, to enable the attention sublayer to jointly attend to information from different representation subspaces at different positions, the attention sublayer employs the multi-head attention shown on the right side of FIG. 2.
[0068] Specifically, to implement the multi-head attention, the attention sublayer applies different attention mechanisms for h in parallel. In other words, the attention sublayer includes different attention layers for h, and each attention layer within the same attention sublayer is configured to receive the same original query Q, original key K, and original value V.
[0069] Each attention layer is configured to transform the original query, key, and value using learned linear transformations and apply the attention mechanism 230 to the transformed query, key, and value. Each attention layer generally learns different transformations from the mutual attention layers within the same attention sublayer.
[0070] Specifically, each attention layer is configured to apply the learned query linear transformation to each original query to generate a layer-specific query for each original query, apply the learned key linear transformation to each original key to generate a layer-specific key for each original query, and apply the learned value linear transformation to each original value to generate a layer-specific value for each original value. Then, the attention layer applies the attention mechanism described above using these layer-specific queries, keys, and values to generate the initial output of the attention layer.
[0071] Next, the attention sub-layer combines the initial outputs of the attention layer to generate the final output of the attention sub-layer. As shown in Figure 2, the attention sub-layer concatenates the outputs of the attention layer and applies a learned linear transformation to the concatenated output to generate the output of the attention sub-layer.
[0072] In some cases, the learned transformation applied by the attention sub-layer reduces the dimensions of the original key, value, and optionally the query. For example, if the dimensions of the original key, value, and query are d and there is an attention layer with h in the sub-layer, the sub-layer may reduce the dimensions of the original key, value, and query to d / h. This maintains the computational cost of the multi-head attention mechanism similar to the computational cost that would be required to run the attention mechanism once with all dimensions, and at the same time increases the representational ability of the attention sub-layer.
[0073] The attention mechanism applied by each attention sub-layer is the same, but the query, key, and value are different for different types of attention. That is, different types of attention sub-layers use different sources for the original query, key, and value received as input by the attention sub-layer.
[0074] Specifically, when the attention sub-layer is an encoder self-attention sub-layer, all of the key, value, and query originate from the same place, in this case the output of a previous sub-network within the encoder, or for the encoder self-attention sub-layer within the first sub-network, the input to the encoder and the embedding of each position can attend to all positions in the order of input. Thus, there are respective keys, values, and queries for each position in the order of input.
[0075] When the attention sublayer is the decoder self-attention sublayer, each position in the decoder attends to all positions in the decoder that precede that position. Thus, all of the keys, values, and queries either arise from the same place, in this case the output of the previous subnetwork of the decoder, or, for the decoder self-attention sublayer within the first decoder subnetwork, from the embeddings of the already-generated outputs. Thus, for each position in output order before the current position, there are respective keys, values, and queries for each position.
[0076] When the attention sublayer is the encoder-decoder attention sublayer, the queries arise from previous components in the decoder, and the keys and values arise from the output of the encoder, i.e., from the encoded representations generated by the encoder. This enables each position in the decoder to attend over all positions in the input sequence. Thus, for each position in output order before the current position, there are respective queries for each position, and for each position in input order, there are respective keys and respective values for each position.
[0077] More specifically, when the attention sublayer is the encoder self-attention sublayer, for each specific input position in input order, the encoder self-attention sublayer is configured to apply an attention mechanism over the encoder subnetwork input at the specific input position using one or more queries derived from the encoder subnetwork input at the specific input position to generate respective outputs for the specific input position.
[0078] When the encoder self-attention sublayer performs multi-head attention, each encoder self-attention layer within the encoder self-attention sublayer applies a learned query linear transformation to each encoder subnetwork input at each input position to generate respective queries for each input position, applies a learned key linear transformation to each encoder subnetwork input at each input position to generate respective keys for each input position, applies a learned value linear transformation to each encoder subnetwork input at each input position to generate respective values for each input position, and then uses the queries, keys, and values to apply an attention mechanism (i.e., the scaled dot-product attention mechanism described above) to determine the initial encoder self-attention output for each input position. The sublayer then combines the initial output of the attention layer as described above.
[0079] When the attention sublayer is a decoder self-attention sublayer, the decoder self-attention sublayer receives inputs for each output position preceding the corresponding output position at each generation time step, and for each of the specific output positions, uses one or more queries derived from the inputs at the specific output position to apply an attention mechanism over the inputs at the output positions preceding the corresponding position to generate an updated representation of the specific output position.
[0080] When the decoder self-attention sublayer performs multi-head attention, each attention layer within the decoder self-attention sublayer applies, at each generation time step, the learned query linear transformation to the inputs at each output position preceding the corresponding output position to generate a respective query for each output position, applies the learned key linear transformation to each input at each output position preceding the corresponding output position to generate a respective key for each output position, applies the learned value linear transformation to each input at each output position preceding the corresponding output position to generate a respective key for each output position, and then uses the queries, keys, and values to apply an attention mechanism (i.e., the scaled dot-product attention mechanism described above) to determine an initial decoder self-attention output for each output position. The sublayer then combines the initial outputs of the attention layers as described above.
[0081] When the attention sublayer is an encoder-decoder attention sublayer, the encoder-decoder attention sublayer receives, at each generation time step, the inputs for each output position preceding the corresponding output position and, for each of the output positions, applies an attention mechanism over the representations encoded at the input positions using one or more queries derived from the input for the output position to generate an updated representation for the output position.
[0082] When an encoder-decoder attention sublayer performs multi-head attention, each attention layer, at each generation time step, applies a learned query linear transformation to the input at each output position preceding the corresponding output position to generate a respective query for each output position, applies a learned key linear transformation to each encoded representation at each input position to generate a respective key for each input position, applies a learned value linear transformation to each encoded representation at each input position to generate a respective value for each input position, and then uses the queries, keys, and values to apply an attention mechanism (i.e., the scaled dot-product attention mechanism described above) to determine an initial encoder-decoder attention output for each input position. The sublayer then combines the initial output of the attention layer as described above.
[0083] FIG. 3 is a flowchart illustrating an exemplary process for generating an output sequence from an input sequence. For convenience, process 300 is described as being executed by one or more computer systems located in one or more locations. For example, a neural network system, such as neural network system 100 of FIG. 1, appropriately programmed in accordance with this specification, can execute process 300.
[0084] The system receives an input sequence (step 310).
[0085] The system processes the input sequence using an encoder neural network to generate an encoded representation of each network input in the input sequence. In particular, the system processes the input sequence through an embedding layer to generate an embedded representation of each network input, and then processes the embedded representation through a sequence of encoder sub-networks to generate an encoded representation of the network input.
[0086] The system processes the encoded representation using a decoder neural network to generate an output sequence (step 330). The decoder neural network is configured to generate an output sequence from the encoded representation in an autoregressive manner. That is, the decoder neural network generates one output from the output sequence at each generation time step. At a given generation time step at which a given output has been generated, the system processes the output preceding the given output in the output sequence through an embedding layer of the decoder to generate an embedded representation. The system then processes the embedded representation through a sequence of decoder subnetwork, linear layer, and softmax layer to generate the given output. Since the decoder subnetwork includes an encoder-decoder attention sublayer and a decoder self-attention sublayer, the decoder utilizes both the outputs that have already been generated and the encoded representation when generating a given output.
[0087] The system can execute process 300 for an input sequence of outputs for which the desired output, i.e., the output sequence to be generated by the system for the input sequence, is unknown.
[0088] The system can also execute Process 300 on a set of training data, i.e., an input sequence of a set of inputs for which the output sequence to be generated by the system is known, in order to train the encoder and decoder to determine values trained on the parameters of the encoder and decoder. Process 300 can be repeatedly executed on inputs selected from the set of training data as part of conventional machine learning training techniques for training initial neural network layers, such as gradient descent by backpropagation training techniques using a conventional optimizer such as the Adam optimizer. During training, the system can incorporate any number of techniques to improve the speed, effectiveness, or both of the training process. For example, the system can use dropout, label smoothing, or both to reduce overfitting. As another example, the system can perform training using a distributed architecture that trains multiple instances of the sequence transformation neural network in parallel.
[0089] This specification uses the term "configured" in relation to system and computer program components. To be configured to perform a particular operation or action by one or more computer systems means that the system has installed thereon software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action during operation. To be configured to perform a particular operation or action by one or more computer programs means that the one or more programs include instructions that cause an apparatus to perform the operation or action when executed by a data processing apparatus.
[0090] The subject matter and the embodiments of the functional operations described in this specification may be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations thereof. Alternatively, or in addition, the program instructions may be encoded in an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information to be transmitted to an appropriate receiver apparatus for execution by a data processing apparatus.
[0091] The term “data processing apparatus” refers to data processing hardware and includes, by way of example, any kind of apparatus, device, and machine for processing data, including a programmable processor, a computer, or multiple processors or computers. The apparatus may also be, or further include, special purpose logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus may optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.
[0092] Such a computer program, which may be referred to or described as a program, software, software application, application, module, software module, script, or code, may be written in any form of programming language, including compiler-type or interpreter-type languages, or declarative or procedural languages, and the computer program may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program may be stored in part of a file that holds other programs or data, such as a markup language document, a single file dedicated to the program, or one or more scripts stored in a plurality of cooperating files, such as a file that stores one or more modules, subprograms, or portions of code. The computer program may be deployed to be executed on one computer located at one site or on a plurality of computers distributed across a plurality of sites and interconnected by a data communication network.
[0093] As used herein, the term "database" is used broadly to denote any collection of data, which need not be structured in any particular way or at all and may be stored on a storage device in one or more locations. Thus, for example, an indexed database can include a plurality of collections of data, each of which may be organized and accessed separately.
[0094] Similarly, as used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers may be dedicated to a particular engine, or in other cases, multiple engines may be installed and executed on the same computer or multiple computers.
[0095] The processes and logical flows described herein may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating input data to produce output. The processing and logical flows may also be performed by special purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0096] A computer suitable for the execution of a computer program may be based on a general purpose or special purpose microprocessor, or both, or any other kind of central processing unit. Generally, the central processing unit receives instructions and data from a read-only memory or a random access memory, or both. Important elements of a computer are a central processing unit for executing or performing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory may be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer also includes, or is operatively coupled to, one or more mass storage devices for storing data, such as, for example, magnetic, magneto-optical disks, or optical disks, or is capable of receiving, or transmitting, or both, data to and from such devices. However, a computer need not have such devices. Further, a computer may be incorporated in another device, such as, for example, a cellular telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.
[0097] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, for example, EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM disks and DVD-ROM disks.
[0098] To provide interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device by which the user can provide input to the computer, such as a mouse or trackball. Other types of devices may be used to provide interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, voice, or tactile input. Additionally, the computer can interact with the user by sending and receiving documents between the device used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from, for example, the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user as a reply.
[0099] A data processing apparatus for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing common and computational entity portions of a machine learning training or production, i.e., inference, workload.
[0100] The machine learning model may be implemented and deployed using a machine learning framework, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.
[0101] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes, for example, backend components as a data server, or includes middleware components such as, for example, an application server, or includes frontend components such as a graphical user interface, a web browser, or an application with which a user can interact with embodiments of the subject matter described herein, or includes any combination of one or more such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication, such as, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs) such as, for example, the Internet.
[0102] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, the server sends data, such as, for example, an HTML page, to a user device for the purpose of displaying data to a user interacting with the device and receiving user input from the user, for example, performing client functions. Data generated at the user device, such as, for example, the result of a user interaction, may be received at the server from the device.
[0103] This specification includes many specific implementation details, but these should not be construed as limiting the scope of any invention or what may be claimed, but rather as an explanation of features that may be specific to a particular implementation of a particular invention. The specific features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable partial combination in multiple embodiments. Moreover, even if a feature is described above as operating in a particular combination and is initially claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to a partial combination or a variation of a partial combination.
[0104] Similarly, operations are shown in the drawings and recited in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or sequence shown, or that all of the recited operations be performed, in order to achieve a desired result. In certain circumstances, multitasking and parallel processing may be beneficial. Moreover, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated into a single software product or packaged into multiple software products.
[0105] Certain embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes shown in the accompanying figures do not necessarily require the particular order or sequence shown to achieve desirable results. In some cases, multitasking and parallel processing may be beneficial.
Description of Reference Numerals
[0106] 100 Neural network system 108 Attention-based sequence conversion neural network 110 Encoder neural network 120 Embedding layer 130 Encoder sub-network 132 Encoder self-attention sub-layer 134 Position-wise feed-forward layer 150 Decoder neural network 152 Output sequence 160 Embedding layer 170 Decoder sub-network 172 Decoder self-attention sub-layer 174 Encoder-decoder attention sub-layer 176 Position-wise feed-forward layer 180 Linear layer 190 Softmax layer 230 Attention mechanism
Claims
1. 1. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to implement a sequence transformation neural network for transforming an input sequence having respective network inputs at each of a plurality of input locations in an input order into an output sequence having respective network outputs at each of a plurality of output locations in an output order, the sequence transformation neural network comprising: an encoder neural network configured to receive the input sequence and generate a respective encoded representation of each of the network inputs in the input sequence, the encoder neural network comprising a sequence of one or more encoder sub-networks, each encoder sub-network configured to receive a respective encoder sub-network input for each of the plurality of input positions and to generate a respective sub-network output for each of the plurality of input positions, each encoder sub-network configured to: an encoder self-attention sublayer configured to receive the sub-network input for each of the plurality of input positions, the encoder self-attention sublayer configured to, for each particular input position in input order: an encoder neural network comprising an encoder self-attention sublayer configured to apply an attention mechanism across the encoder sub-network inputs at the particular input location using one or more queries derived from the encoder sub-network inputs at the particular input location to generate a respective output for the particular input location; a decoder neural network configured to receive the encoded representation and to generate the output sequence.
2. The encoder neural network includes: For each network input in the input sequence: mapping the network inputs to an embedding representation of the network inputs; combining the embedded representations of the network inputs with positional embeddings of the input positions of the network inputs in input order to generate a combined embedded representation of the network inputs; 2. The system of claim 1 , further comprising an embedding layer configured to provide the combined embedded representation of the network inputs as the encoder sub-network input to a first encoder sub-network in the sequence of encoder sub-networks.
3. 3. The system of claim 1 or 2, wherein the respective encoded representations of the network inputs are the encoder sub-network outputs produced by the last encoder sub-network in the sequence.
4. 4. The system of claim 1, wherein for each encoder subnetwork other than the first encoder subnetwork in the sequence, the encoder subnetwork input is the encoder subnetwork output of the preceding encoder subnetwork in the sequence.
5. At least one of the encoder sub-networks For each input position, receiving an input at the input location; 5. The system of claim 1, further comprising a positional feedforward layer configured to apply a sequence of transformations to the input at the input location to generate an output for the input location.
6. 6. The system of claim 5, wherein the sequence comprises two learned linear transformations separated by an activation function.
7. The at least one encoder sub-network a residual connection layer that combines the outputs of the per-position feedforward layer with the inputs to the per-position feedforward layer to generate a residual output for each encoder position; and a layer normalization layer that applies layer normalization to a residual output for each encoder position.
8. Each encoder sub-network is a residual connection layer that combines the outputs of the encoder self-attention sublayer with the inputs to the encoder self-attention sublayer to generate an encoder self-attention residual output; and a layer normalization layer that applies layer normalization to the encoder self-attention residual output.
9. 9. The system of claim 1, wherein each encoder self-attention sub-layer comprises a plurality of encoder self-attention layers.
10. Each encoder self-attention layer applying the learned query linear transformation to each encoder sub-network input at each input location to generate a respective query for each input location; applying the learned key linear transformation to each encoder sub-network input at each input position to generate a respective key for each input position; configured to apply the learned value-linear transformation to each encoder sub-network input at each input location to generate a respective value for each input location; For each input position, determining an input-position-specific weight for each of said input positions by applying a comparison function between said query and said key for said input position; 10. The system of claim 9, configured to determine an initial encoder self-attention output for the input position by determining a weighted sum of values of the input position weighted by a weight specific to the corresponding input position.
11. 11. The system of claim 10, wherein the encoder self-attention sublayer is configured to, for each input position, combine the initial encoder self-attention outputs for the input position generated by the encoder self-attention sublayer to generate an output of the encoder self-attention sublayer.
12. 12. The system of claim 9, wherein the encoder self-attention layers operate in parallel.
13. 13. The system of claim 1, wherein the decoder neural network generates the output sequence in an autoregressive manner by generating, at each of a plurality of generation time steps, a network output at a corresponding output position conditioned on the encoded representation and on a network output at an output position preceding the output position in the output order.
14. 14. The system of claim 13, wherein the decoder neural network comprises a sequence of decoder sub-networks, each decoder sub-network configured to receive, at each generation time step, a respective decoder sub-network input for each of the plurality of output positions preceding the corresponding output position, and to generate a respective decoder sub-network output for each of the plurality of output positions preceding the corresponding output position.
15. The decoder neural network comprises: At each generation time step for each network output at an output position preceding said output position in said output order: mapping the network output to an embedding representation of the network output; combining the embedded representations of the network outputs with positional embeddings of the output positions of the network outputs in output order to generate a combined embedded representation of the network outputs; 15. The system of claim 14, further comprising an embedding layer configured to provide the combined embedded representation of the network output as an input to a first decoder sub-network in the sequence of decoder sub-networks.
16. At least one of the decoder sub-networks for generating an output for the output location; At each generation time step For each output position preceding the corresponding output position, receiving an input at the output location; 16. The system of claim 14 or 15, comprising a position-wise feedforward layer configured to apply a sequence of transformations to the input at the output positions.
17. 17. The system of claim 16, wherein the sequence comprises two learned linear transformations separated by an activation function.
18. The at least one decoder sub-network a residual connection layer that combines the outputs of the per-position feedforward layer with the inputs to the per-position feedforward layer to generate a residual output; 18. The system of claim 16 or 17, further comprising a layer normalization layer that applies layer normalization to the residual output.
19. Each decoder sub-network is At each generation time step, receiving an input for each output location preceding said corresponding output location; and for each of said output locations:
14. The system of claim 10, further comprising an encoder-decoder attention sublayer configured to apply an attention mechanism across the encoded representation at the input location using one or more queries derived from the input at the output location to generate an updated representation for the output location.
20. Each encoder-decoder attention sub-layer comprises a plurality of encoder-decoder attention layers, each of which, at each generation time step, applying the learned query linear transformation to the input at each output position preceding the corresponding output position to generate a respective query for each output position; applying the learned key linear transformation to each encoded representation at each input position to generate a respective key for each input position; applying the learned value-linear transformation to each encoded representation at each input location to generate a respective value for each input location; For each output position preceding the corresponding output position, determining an input-position-specific weight for each of said input positions by applying a comparison function between said query and said key of said output position; 16. The system of claim 15, configured to determine an initial encoder-decoder attention output for the output position by determining a weighted sum of values of the input positions weighted by the corresponding output position-specific weights.
21. 21. The system of claim 20, wherein the encoder-decoder attention sublayer is configured to combine the encoder-decoder attention outputs generated by the encoder-decoder layer at each generation time step to generate the output of the encoder-decoder attention sublayer.
22. 22. The system of claim 20 or 21, wherein the encoder-decoder attention layers operate in parallel.
23. Each decoder sub-network is a residual connection layer that combines the outputs of the encoder-decoder attention sublayer with the inputs to the encoder-decoder attention sublayer to generate a residual output; 23. The system of claim 19, further comprising a layer normalization layer that applies layer normalization to the residual output.
24. Each decoder sub-network is At each generation time step, receiving an input for each output location preceding said corresponding output location; and for each of said particular output locations:
24. The system of claim 14, further comprising a decoder self-attention sublayer configured to apply an attention mechanism across the inputs at the output positions preceding the corresponding position using one or more queries derived from the inputs at the particular output position to generate an updated representation for the particular output position.
25. Each decoder self-attention sub-layer comprises a plurality of decoder self-attention layers, each decoder self-attention layer comprising, at each generation time step: applying the learned query linear transformation to the input at each output position preceding the corresponding output position to generate a respective query for each output position; applying the learned key linear transformation to each input at each output position preceding the corresponding output position to generate a respective key for each output position; applying the learned value-linear transformation to each input at each output position preceding the corresponding output position to generate a respective key for each output position; For each output position preceding the corresponding output position, determining a respective input location specific weight for each of said output locations by applying a comparison function between said query and said key for said output location; 25. The system of claim 24, configured to determine an initial decoder attention output for the output position by determining a weighted sum of the values weighted by the corresponding output position specific weight of the output position.
26. 26. The system of claim 25, wherein the encoder-decoder attention sublayer is configured to combine the encoder-decoder attention outputs generated by the encoder-decoder layer at each generation time step to generate the output of the encoder-decoder attention sublayer.
27. 27. The system of claim 25 or 26, wherein the encoder-decoder attention layers operate in parallel.
28. Each decoder sub-network is a residual connection layer that combines the outputs of the decoder self-attention sublayer with the inputs to the decoder self-attention sublayer to generate a residual output; 28. The system of claim 24, further comprising a layer normalization layer that applies layer normalization to the residual output.
29. 29. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to implement the sequence transformation neural network of any one of claims 1 to 28.
30. receiving an input sequence having a respective input at each of a plurality of input locations in an input order; processing the input sequence through the encoder neural network of any one of claims 1 to 28 to generate respective encoded representations of each of the inputs of the input sequence; and processing the encoded representation through the decoder neural network of any one of claims 1 to 28 to generate an output sequence having a respective output at each of a plurality of output positions in output order.
31. 31. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method of claim 30.
32. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations of the method recited in claim 31.
Citation Information
Patent Citations
Unsupervised matching in fine-grained datasets for single-view object reconstruction
US20170124433A1
Generating target sequences from input sequences using partial conditioning
US20170140753A1