Generating continuous valued data with a transformer neural network

By employing a Transformer neural network with attention layers and a continuous distribution output block, the method addresses the challenges of generating continuous valued data, enhancing performance and efficiency in data generation tasks.

WO2025104345A1PCT designated stage expired Publication Date: 2025-05-22DEEPMIND TECH LTD

Patent Information

Application Number
PCT/EP2024/082740
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-01
Filing Date
2024-11-18
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing neural networks face challenges in generating continuous valued data, particularly due to the need for discrete tokens and fixed vocabularies, which can lead to training difficulties and inefficiencies.

Method used

The use of a Transformer neural network with a sequence of attention neural network layers and an output block that processes attention layer outputs to generate parameters defining a continuous distribution over possible values of a data element, allowing for the sampling of continuous valued data elements.

Benefits of technology

This approach enables the generation of high-quality continuous valued data without the limitations of discrete tokens, improving performance and reducing computational and memory costs, particularly in tasks like image generation and dense prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082740_22052025_PF_FP_ABST
    Figure EP2024082740_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods, implemented as computer programs on one or more computers in one or more locations, for generating a sequence of data elements using a neural network comprising a sequence of attention neural network layers. The sequence comprises a respective continuous valued data element at each position in a sequence of positions. Implementations of the described techniques remove the need for discrete tokens and fixed, finite vocabularies.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATING CONTINUOUS VALUED DATA WITH A TRANSFORMER NEURAL NETWORKCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Application No. 63 / 600,433, filed on November 17, 2023, to U.S. Provisional Application No. 63 / 565,204, filed on March 14, 2024, and to U.S. Provisional Application No. 63 / 702,090, filed on October 1, 2024. The disclosure of the prior applications is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND

[0002] This specification relates to neural networks.

[0003] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters, e.g. weights.SUMMARY

[0004] This specification describes systems and methods, implemented as computer programs on one or more computers in one or more locations, for generating a sequence of data elements using a neural network comprising a sequence of attention neural network layers. The sequence comprises a respective continuous valued data element at each position in a sequence of positions. Implementations of the described techniques remove the need for discrete tokens and fixed, finite vocabularies.

[0005] In one aspect there is described a computer-implemented method of generating a sequence of data elements that has a respective continuous valued data element at each position in a sequence of positions. The method involves, e.g. for each position after a first position in the sequence of positions, obtaining a current sequence of data element embeddings that comprises a respective data element embedding of (at least) each data element at a position that precedes the current position. Each data element embedding represents a value of the data element that is a continuous variable.

[0006] The method involves processing the current sequence of data element embeddings using a neural network to generate the data element at (at least) the current position. The neural network comprises a sequence of attention neural network layers and an output block, wherein each attention layer has an attention layer input for each element of the sequence of data elements and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output. The output block performs operations comprising processing the attention layer output of a last attention neural network layer before the output block to generate parameters defining a continuous distribution over possible values of a data element, and sampling from the continuous distribution to obtain the value of the data element at (at least) the current position, wherein the value of the data element at the current position is a continuous variable.

[0007] In some implementations performing the method for each position after a first position in the sequence of positions involves obtaining the current sequence of data element embeddings comprising the respective data element embedding of (just) each data element at a position, i.e. a respective position, that precedes the current position. The current sequence of data element embeddings is processed using the neural network to generate the data element at the current position, and the output block performs operations comprising sampling from the continuous distribution to obtain the value of the data element at the current position.

[0008] In some implementations performing the method for each position after a first position in the sequence of positions involves performing the method in parallel. That is, the method can involve obtaining the current sequence of data element embeddings comprising the respective data element embedding of each data element in the sequence of data elements. The current sequence of data element embeddings can be processed in parallel using the neural network to generate each data element in the sequence of data elements. The output block can perform operations comprising processing the attention layer output of a last attention neural network layer before the output block to generate, for each data element in parallel, parameters defining the continuous distribution over possible values of the data element, and sampling from the continuous distribution to obtain the value of each data element.

[0009] There is also described a computer-implemented method of training a neural network for generating a sequence of continuous valued data elements. An example of this method involves obtaining an autoencoder, e.g. a trained autoencoder, comprisingan encoder neural network and a decoder neural network. The neural network can then be trained, for a plurality of the training data items, by processing the training data item using the encoder neural network to generate the first set of parameters defining the posterior distribution of the set of latent variables. The posterior distribution is sampled from to determine training sampled values of the set of latent variables, and the neural network is trained using the training sampled values of the set of latent variables.

[0010] There is also described a computer-implemented method of generating image data, the image data specifying values for pixels of an image. The method involves generating a sequence of data elements that comprises a respective continuous valued data element at each position in a sequence of positions. The generation comprises, for each position after a first position in the sequence of positions, obtaining a current sequence of data element embeddings that comprises a respective data element embedding of each data element at a position that precedes the current position, wherein each data element embedding represents a value of the data element that is a continuous variable. The generation further comprises processing the current sequence of data element embeddings using a neural network to generate the data element at the current position. The neural network comprises a sequence of self-attention neural network layers and an output block, wherein each self-attention layer has an attention layer input for each element of the sequence of data elements and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output. The output block performs operations comprising processing the attention layer output of a last self-attention neural network layer before the output block to generate parameters defining a continuous distribution over possible values of a data element, sampling from the continuous distribution to obtain the value of the data element at the current position, wherein the value of the data element at the current position is a continuous variable and processing the sequence of data elements using a trained decoder neural network to generate the image data.

[0011] There is also described a computer-implemented method of training a neural network using a sequence of continuous-valued data elements. The method involves processing each data element using a normalizing flow model to generate a respective training data item (each training data item comprising a vector with a dimensionality matching a dimensionality of the respective continuous-valued data element). In this way a sequence of training data items can be obtained. The neural network can be trained to maximize a likelihood of a next element in the sequence of training dataitems. In implementations the neural network can be trained jointly with the normalizing flow model, end-to-end, e.g. by backpropagating gradients of a training objective through both the neural network and the normalizing flow model.

[0012] There is also described a computer-implemented method of generating values for pixels of an image. The method involves autoregressively extending a sequence of soft tokens. Each soft token comprises a vector of continuous values that represents a patch of a sequence of patches of the image, e.g. that tile the image. Each soft token of the sequence of soft tokens can be processed using a trained invertible model to generate a set of pixel values for each respective patch of the image.

[0013] This specification also provides one or more computers, and one or more storage devices or computer storage media, storing instructions that, when executed by the one or more computers, cause the one or more computers to perform any of the methods disclosed herein.

[0014] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.

[0015] Standard discrete-token generative transformer models take discrete tokens and use a lookup table to determine an input embedding to be processed by a discrete-token generative transformer. The output from the discrete-token generative transformer is a categorical distribution over a finite vocabulary. Such discrete-token generative transformers may be trained using a vector-quantized variational autoencoder (VQ- VAE). In contrast to the standard discrete-token generative transformer model, the neural network described herein (which could be a transformer neural network) takes as input a sequence of continuous-valued, i.e. real valued, data elements and predicts the parameters of a continuous distribution.

[0016] It has been found that using continuous-valued data (e.g. data comprising real numbers, and not formed from a discrete set) improves upon prior methods which use discrete valued data. For example, using continuous-valued data side-steps training difficulties such as low codebook usage in VQ-VAEs and corresponding mitigations like entropy losses or codebook-splitting algorithms, by enabling the use of standard VAEs which are much easier to train. Furthermore, the techniques disclosed herein avoid large embedding matrices because the feature representations can directly be consumed and predicted by the neural network. Such matrices can consume large amounts of memory, particularly in multimodal settings, and the described techniques can avoid this and are also straightforward to use with interleaved multimodal tokens.The quantization-free approach described herein outperforms its VQ-based equivalent when using causal transformers for class-conditional image generation, often by a large margin, and can also improve on Mask-GIT -based approaches. The describe techniques also perform well on dense prediction tasks such as panoptic segmentation and monocular depth prediction.

[0017] Implementations of the system combine a neural network, e.g. a Transformer neural network, with an invertible model, in particular a normalized flow model. The normalized flow model facilitates training jointly with the neural network, and can generate data items with better detail, particularly for out of distribution data items, such as images containing text. Such an architecture is also more general. The normalized flow model can be trained jointly with the neural network, end-to-end, to encourage the normalized flow model to leam embeddings that facilitate their modelling with the neural network, e.g. autoregressive modelling using the Transformer neural network.

[0018] Broadly the described techniques can provide improved performance, e.g. better image generation quality, at significantly lower computational cost, and with lower memory requirements.

[0019] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] FIG. 1 shows an example system for generating a sequence of data elements.

[0021] FIG. 2 shows an example encoder.

[0022] FIGS. 3 A and 3B illustrate an example of a technique for training a neural network and an example of a technique for generating an image using a neural network.

[0023] FIG. 4 is a flow diagram of an example process for generating a sequence of data elements.

[0024] FIG. 5 is a flow diagram of an example process for generating image data.

[0025] FIG. 6 is a flow diagram of an example process for training a neural network to generate a sequence of data elements.

[0026] FIG. 7 illustrates a process for training a neural network combined with a normalizing flow model.

[0027] FIG. 8 illustrates a process for using a trained neural network combined with a trained normalizing flow model for generating an image in response to a prompt.

[0028] FIG. 9 is a flow diagram of an example process for generating a sequence of continuous-valued data elements using a neural network combined with an invertible model.

[0029] FIGS. 10 A- 10C show example data items generated by the described techniques.

[0030] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0031] Broadly this specification describes techniques for using a neural network to generate vector sequences with real -valued entries, i.e. a sequence of data elements with continuous valued data element at each position in the sequence.

[0032] In implementations the neural network comprises a sequence of attention neural network layers that may be, but need not be, self-attention neural network layers; and an output block. Each attention layer has an attention layer input for each element of the sequence of data elements and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output.

[0033] In some implementations the neural network comprises a Transformer neural network. Such a Transformer neural network can be a so-called decoder(-only) Transformer, configured to process a current input sequence of data element embeddings for generating the next (current) data element, in particular generating elements of the sequence sequentially, in an autoregressive manner. Or such a Transformer neural network can be a so-called encoder Transformer, configured to process an input sequence of data element embeddings for generating an output sequence in parallel. In some implementations the neural network comprises a so-called encoder-decoder Transformer neural network, in particular an encoder Transformer neural network that provides an encoder Transformer neural network output comprising features that are used by the decoder Transformer neural network, e.g. by cross-attending to these.

[0034] In general a Transformer neural network can be characterized by having a succession of attention, e.g. self-attention, neural network layers. Each attention neural network layer has an attention layer input for each element of the input and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output for each element of the input. In implementations the attention mechanism computes a similarity between a query and a set of key -value pairs. In implementations one or both (in the case of self-attention) of the query and the set of key -value pairs are determined from the attention layer input.

[0035] As an example, in a self-attention neural network layer an input embedding may be used to determine a query vector and a set of key -value vector pairs, that are used to generate an updated embedding comprising a weighted sum of the values, weighted by a similarity function of the query to each respective key. The similarity function may comprise, e.g., a dot product, cosine similarity, or other similarity measure; the query, keys, and values may all be vectors. For example the attention mechanism may be configured to apply each of a query transformation e.g. defined by a matrix WQ, a key transformation e.g. defined by a matrix WK, and a value transformation e.g. defined by a matrix Wv, to the attention layer input for each element of an input sequence X to derive a respective query vector Q = XWQ, key vector K = XWK, and value vector V = XWVwhich are used determine an attended sequence for the output.

[0036] In general the sequence of attention neural network layers, e.g. the Transformer neural network, can process the input sequence of data element embeddings, e.g. a sequence of vectors, to generate an attention layer output that also comprises a sequence of embeddings, e.g. vectors, each corresponding to one of the input sequence of vectors. When used in an autoregressive manner just the last, i.e. most recently, generated embedding, e.g. from the “decoder” Transformer, can be used to obtain next the element of the sequence of data elements. Alternatively the sequence of attention neural network layers, can be used to process the input sequence of data element embeddings to generate an output sequence of embeddings in parallel, e.g. when configured as an “encoder” Transformer neural network. For example an encoder Transformer type architecture as described herein can be used in place of the conventional Transformer in Mask-GIT (Chang et al„ “Mask-GIT: Masked Generative Image Transformer”, arXiv:2202.04200vl, Feb 2022).

[0037] Throughout this specification, an “embedding” of an entity can refer to a representation of the entity as an ordered collection of numerical values, e.g., a vector or matrix of numerical values. An embedding of an entity can be generated, e.g., as the output of a neural network that processes data characterizing the entity, or as a result of some other encoding process.

[0038] In general the neural network outputs described herein may define distributions, e.g. as a set of scores or by outputting parameters defining a distribution. A particular value may be obtained by sampling from such a distribution or the particular value may be defined e.g. by a mean or maximum value of the distribution.

[0039] In implementations the neural network, e.g. the Transformer neural network, comprises an output block that performs operations that involve processing the attention layer output of a last attention neural network layer before the output block. The output block can generate parameters defining a continuous distribution over possible values of a data element. A sample can be obtained from the continuous distribution to obtain the value of the data element at the current position. The value of the sampled data element, e.g. at the current position, is a continuous variable.

[0040] The data elements generated and processed by the neural network can be referred to as tokens. In the described system such tokens have continuous values rather than being selected from a discrete vocabulary, and can be referred to as “soft tokens”. The described neural network can be referred to as a Generative Infinite-Vocabulary Transformer (GIVT).

[0041] As used herein a soft token can comprise an embedding, i.e. an ordered collection of numerical values, in particular a vector with continuous-valued elements. That is, a soft token represents a token but is not restricted to discrete values that represent tokens in a predetermined vocabulary of tokens, i.e. a soft token does not represent any particular item in a vocabulary such as a word or sub-word. A soft token can be an embedding in a learned data item embedding space of the system.

[0042] As described further later, the data elements or soft tokens that are generated can directly represent e.g. values of pixels of an image or of an audio waveform, or can represent values of latent variables, e.g. of a trained autoencoder, that are decoded to generate, e.g. pixels of an image or an audio waveform. In general the described techniques can be used for any form of image, audio, or time-series modelling; they can also be used for image processing as described later.

[0043] In one example implementation the described neural network (GIVT) can be obtained by modifying a Transformer neural network to, at the input, replace a finite- vocabulary lookup table with a linear projection of the input vectors, and at the output replace a logits prediction mapped to a categorical distribution with a prediction of the parameters of a continuous distribution, e.g. a multivariate Gaussian mixture model.

[0044] In some implementations where a distribution to be modelled by the generated data elements is complex, e.g. a distribution of latent variables, the distribution can be preprocessed using a small invertible (normalizing) flow model; this is also referred to herein as an “adapter”. Broadly such an adapter implements a learned mapping that maps every input maps to a unique output and vice versa, i.e. the mapping is invertible(bijective). The adapter, which is jointly trained with the neural network (no additional losses needed), can facilitate matching the distribution to be modelled to that of the generated data elements. In inference samples from the neural network (GIVT) are processed by the inverted adapter.

[0045] A neural network implementation of such an adapter which is “volume preserving” can involve dividing an input into two parts, apply a learned transformation to only one part, using one or more neural network layers, whilst leaving the other unchanged, combining these to obtain the output. The transformation is invertible (bijective) because the input can be recovered by dividing the output into two parts and applying the same transformation to one part of the output whilst leaving the other unchanged, combining these to obtain the input (see, e.g. Dinh et al., “NICE”, arXiv: 1410.8516v6, Apr 2015, at section 3.2; and Dinh et al. arXiv: 1605.08803v3, Feb 2017).

[0046] In some implementations, again where a distribution to be modelled by the generated data elements is complex, e.g. where the data elements directly model, say pixels of an image or audio waveform values, the data element embeddings can be processed using a normalizing flow model to generate a corresponding sequence of processed elements, e.g. that represent the pixels of the image or audio waveform values. This normalizing flow model is typically larger and more powerful than the previously mentioned “adapter”.

[0047] Examples of normalizing flow models are described in e.g. Rezende et al., “Variational Inference with Normalizing Flows”, arXiv: 1505.05770, 2016; Kingma and Dhariwal, “Glow: Generative Flow with Invertible 1 >< 1 Convolutions”, 2018; and elsewhere. In general this also transforms a probability distribution using an invertible (bijective) mapping and can transform between simple, e.g. uniform of Gaussian distributions, and complex distributions. Depending on the transform used these may, but do not necessarily preserving the total probability mass (density), in which case an additional objective function term can be used that comprises the log of the absolute value of the determinant of the Jacobian of the transformation (arising from a change in volume in the input space due to the transform).

[0048] The data element embedding for the first position in the sequence may comprise an embedding for a start of sequence token and / or an embedding of conditioning data. Additionally, or alternatively conditioning data may be concatenated to the embedding of each element of the sequence. Additionally, or alternatively the conditioning datamay be provided as a side input to the neural network, i.e. separately from the data element embeddings.

[0049] Generally, “conditioning” a neural network can refer to providing conditioning data to the neural network as an input, such that the neural network jointly processes the conditioning data together with any other neural network inputs, e.g., the current sequence of data element embeddings. Generally, the conditioning data can be processed by the neural network in any appropriate manner. As one example the conditioning data may be provided as one of the data elements, e.g. a data element on the first position in the sequence of positions. As another example it can be processed as an additional input by one or more of the self-attention blocks. A few examples of conditioning data are described later.

[0050] As previously described, in some implementations the system autoregressively generates a sequence of data elements, i.e., such that each data element in the sequence is conditionally generated based on the preceding data elements in the sequence. Then the attention applied by the attention neural network layers may be causally masked, so that it is only over earlier embeddings in the sequence.

[0051] Referring now to FIG. 1, this shows a system 100 implemented as computer programs on one or more computers in one or more locations for generating a sequence of data elements. The system of FIG. 1 can be implemented as computer programs on one or more computers in one or more locations.

[0052] The system 100 comprises a neural network 102. The neural network comprises a sequence of attention layers 110, which may be, but need not be, self-attention layers; and an output block 112.

[0053] An attention layer is attention layer can be one that is configured to implement an attention mechanism. The attention mechanism can map a query and a set of key -value pairs, each derived from the attention layer input to an output from which the attention layer output is derived. The output can be computed as a weighted sum of the values, where the values are weighted by a similarity function of the query to each respective key. In a self-attention layer the query and the set of key-value pairs are each derived from the attention layer input.

[0054] In implementations the neural network 102 includes an additional linear neural network layer, i.e. a layer without a non-linear activation function, at its input (not shown in FIG. 1). The additional linear neural network layer can process an input to the neuralnetwork 102 and provide an output to a first attention neural network layer of the sequence of attention neural network layers of the neural network.

[0055] The output block 112 can be configured to process the attention layer output of a last attention neural network layer (of the sequence of attention layers 110 before the output block) to generate parameters of a continuous distribution over possible values of a data element.

[0056] In implementations the neural network 102 is configured to process continuous valued data elements, i.e. the continuous valued data elements can comprise d- dimensional real numbers, and can generate, i.e. parameterize, a continuous distribution.

[0057] The neural network can generate a sequence of data element embeddings by processing a current sequence of data elements using the linear layer; i.e. using a linear projection of the input vectors. The current sequence of data element embeddings can provide the attention layer input to the first attention, e.g. self-attention, neural network layer of the sequence of attention neural network layers of the neural network.

[0058] The output block 112 performs can generate parameters defining a single Gaussian distribution, or a multivariate Gaussian distribution (assuming a diagonal covariance matrix or with a full covariance matrix). For example parameters defining the continuous distribution generated by the neural network 102 may comprise parameters of a Gaussian mixture model, e.g. a weight (or “mixing coefficient”), mean, and variance for Gaussian distribution of the mixture. As an example a ^-component Gaussian mixture model for a value, x, with component weights TT for multivariate Gaussian distributions JV'(x|JU Sj) with meanand, e.g. diagonal, covariance matrix

[0059] Processing a sequence of data element embeddings, e.g. the current sequence of data element embeddings using the neural network 102, as described herein, can also involve adding a position encoding to each data element embedding, e.g. a learned position encoding, a sinusoidal position encoding, or a rotary position encoding.

[0060] Processing the current sequence of data element embeddings using the neural network 102 to generate the data element at the current position can involve processing the current sequence of data element embeddings with the neural network conditioned on a conditioning input that defines one or more characteristics of the generated sequence of data elements.

[0061] As one example, where the neural network 102 includes one or more crossattention neural network layers the conditioning input may be incorporated using crossattention to the conditioning input. As another example the conditioning input may be included in one of the data elements, e.g. as one or more embeddings prepended to the current sequence of data element embeddings (as a prompt). For example, for classconditional image generation a [CLS] vector, i.e., a learned vector for each class c, can be prepended to the input sequence.

[0062] In some implementations the system 100 includes a decoder neural network 106, e.g. where the neural network 102 is configured to generate a set or sequence of latent variables z. The decoder neural network 106 process the set of latent variables, z, to generate the sequence of data elements that together can constitute a data item such as an image or audio data item.

[0063] The decoder neural network 106 can have been pre-trained, e.g. in combination an encoder neural network 104, shown in FIG. 2, that together with the decoder neural network 106 forms an autoencoder e.g., a variational autoencoder (VAE). Alternatively the decoder neural network 106 and the encoder neural network 104 can be trained jointly with the neural network 102, as described later.

[0064] In implementations the encoder neural network 104 is configured to process an encoder input, e.g. a data item of the type generated by the decoder neural network 106, to generate a set of latent variables that encodes the data item. For example the encoder neural network 104 can generate a first set of parameters defining a posterior distribution of the set of latent variables from which the set of latent variables can be sampled.

[0065] In general the encoder neural network 104 and the decoder neural network 106 may have any suitable architecture and may include, e.g., one or more feed forward neural network layers, one or more recurrent neural network layers, one or more convolutional neural network layers, one or more attention neural network layers, or one or more normalization layers.

[0066] In some implementations, described further later, the system 100 does not include the decoder neural network 106, or the encoder neural network 104. Instead the neural network 102 can directly generate a sequence of data elements that together constitute a data item such as an image or audio data item.

[0067] As a particular example, the neural network 102 can be arranged to generate a sequence of vectors zi, Z2, ... zn, with real-valued entries, rather than discrete tokens as typically used in other transformer models. For example, the neural network 102 can inputa current sequence of data element embeddings zi, z , ..., zn-i and generates a data element at the current position, zn. The neural network 102 can be arranged to repeat this process, e.g. until a defined number of data elements has been generated or until a value indicating termination of the sequence is generated.

[0068] When all the data elements have been generated, these may be output to the decoder neural network 106 as a set of latent variables z. The decoder neural network 106 decodes the set of latent variables and outputs a decoded output, which may be an image, data defining an audio signal waveform, or another type of data item.

[0069] In some implementations the encoder 104, shown in FIG. 2, forms part of a variational autoencoder. The encoder neural network 104 can then be configured to process an encoder input, e.g. an image, to generate a first set of parameters, / / , , defining a posterior distribution of a set of latent variables. The posterior distribution can be sampled to determine sampled values of the set of latent variables. During training of the neural network 102, as discussed in more detail with respect to FIG. 3 A, the neural network 102 can be trained to model the set of latent variables the encoder neural network 104 as a sequence of data elements.

[0070] FIG. 3 A is a schematic representation of a specific and purely illustrative example process for training the neural network 102. In this illustration the latent variables are generated by the encoder neural network 104 of a trained variational autoencoder (VAE), but other encoders can be used or the encoder, and decoder, can be omitted entirely.

[0071] In the particular example of FIG. 3 A the VAE may be a so-called P-VAE, in which the regularization (KL divergence) term is multiplied by a factor ft (Higgins et al., “Beta- VAE: Learning basic visual concepts with a constrained variational framework”, ICLR 2016), hyperparameter fi controlling how strongly the set of latent variables is regularized. This allows adjustment of the information in the latents and hence adjustment of a balance between the neural network 102 and the decoder 106, e.g. increasing ft reduces the information in the latents, simplifying modelling by the neural network 102 (and shifting reconstruction effort to the decoder).

[0072] In FIG. 3 A the encoder neural network 104 generates a first set of parameters , a, defining a posterior distribution of a set of latent variables. The set of latent variables are real valued, e.g. each represents a value of the data element that is a continuous variable. That is, given an input image, the encoder 306 predicts mean / / , and covariance a of a multivariate normal distribution with diagonal covariance matrix. The posterior distribution is sampled to determine one or more sampled values of the set of latentvariables z. That is, a sequence of real-valued latent vectors is sampled from the encoder 104 output.

[0073] In this example, where conditioning is used, conditioning data c can be added to the sampled values to obtain the set of latent variables z. For example, the conditioning data c may be concatenated with the sampled values from the encoder neural network 104. The conditioning data c may define a characteristic of the training data item, such as indicating a category of a set of categories in which the training data item belongs.

[0074] In the example of FIG. 3 A, the encoder neural network 104 includes a subsystem to spatially-down-sample an H W image training data item to generate an h*w set of latent variables z (e.g. with h = [77 / 16], w=[IF / 16]), each comprising a -dimensional feature vector with associated and a values. The KL divergence term can be computed for the whole image, e.g. by flattening the associated / / and a with shapes w * h x d into whd vectors.

[0075] The set of latent variables z is provided to the neural network 102, which is trained with to model the continuous-valued latent sequences from encoder 104. That is, the neural network 102 is trained to predict (z), or (z|c) when a conditioning signal, c, is used, e.g. in class-conditional generation. Where the neural network 102 autoregressively generates the sequence of data elements the neural network 102 can be trained using teacher forcing, i.e. using ground truth rather than predicted (partial) input sequences.

[0076] As shown in FIG. 3 A, the set of latent variables z can be reshaped into an hw- length sequence of -dimensional real-valued vectors that are processed by the neural network 102. The -dimensional real-valued vectors are “soft tokens” rather than discrete-values tokens (and no embedding lookup tables are required at the input of the neural network 102). As previously described, in some implementations a single linear layer may be used to project from d to the Transformer hidden dimension, i.e. the hidden dimension of the first attention neural network layer of the neural network 102.

[0077] At the output, the neural network 102 does not predict a categorical distribution, and instead predicts parameters of a continuous distribution. The continuous distribution can be modelled with a ^-mixture Gaussian mixture model (GMM); merely as an example, 2 < k < 32. The neural network 102 can then predict 2kd + k parameters per soft token (kd mean and kd variance parameters for the k mixture components, and k mixture probabilities, i.e. weights). It can be beneficial to normalize the mixture probabilities with a softmax activation, and the variance parameters with softplus.

[0078] The neural network 102 can be trained using a cross-entropy loss (equivalent to minimizing the negative log-likelihood, NLL) on the distribution p predicted by the neural network 102, e.g. to minimize the loss L =cEz[-- logp (z|c)], using the conditioning signal c when available, e.g., in class-conditional generation; and assuming that the classes or conditioning signal c is uniformly distributed.

[0079] Continuing the above particular example, in which the neural network 102 generates the sequence of data elements autoregressively, the neural network 102 can be trained to predict every -dimensional vector in the hw sequence of latents conditioned on all previous vectors. The (self-)attention layers can then be causally masked, i.e. to prevent positions in the sequence from attending to future positions in the sequence. As previously described, for class-conditional image generation a [CLS] vector can prepended to the input sequence, i.e., a learned vector for each class c.

[0080] Where the neural network 102 generates the (complete) sequence of data elements in parallel from the input sequence of data element embeddings, e.g. where the neural network 102 implements an encoder Transformer, the causal masking is omitted. The neural network 102 can be trained by randomly masking tokens (data element embeddings) so that the neural network 102 predicts these tokens. The neural network 102 can be trained using any suitable prediction loss, e.g. a cross-entropy loss.

[0081] Referring again to FIGS. 1 and 2, the encoder neural network 104 and decoder neural network 106 can be pre-trained, or the neural network 102 can be trained, or finetuned, end-to end with the encoder neural network 104, e.g. based on any suitable reconstruction loss, such as an MSE loss. In general training involves backpropagating gradients of an objective function, e.g. loss, to adjust trainable parameters (such as weights) of the encoder neural network 104, decoder neural network 106, or neural network 102.

[0082] Where the encoder neural network 104 and decoder neural network 106 together comprise a variational autoencoder (VAE) that encodes and subsequently decodes (reconstructs) a data item x, the VAE can be pre-trained using an objective function that comprises a reconstruction or likelihood term (loss), such as an MSE (mean squared error) term or, for image generation a perceptual loss or a GAN loss for image generation. More particularly the objective function can comprise a sum of the reconstruction or likelihood term and a regularization term, e.g. a KL divergence term. This term can be estimated for a one or a batch of training data item or, where a Gaussian encoder distribution is used, can be calculated in closed form, e.g. as described in Kingma and Welling,arXiv: 1312.6114, 2013, at F.l. Backpropagation through the step of sampling from the posterior distribution ( / / ; a) can be avoided using the re-parametrization trick (Kingma and Welling, ibid), which replaces the sampling by added noise.

[0083] FIG. 3B is a schematic representation of inference with the neural network 102. During inference the neural network 102 is sampled from either sequentially (as shown in FIG. 1 and FIG. 3B), or in parallel, to obtain a sampled sequence of data elements. In the example of FIG. 1 and FIG. 3B the sequence of data elements is generated autoregressively and defines values of a set of latent variables, z, which is decoded using the decoder neural network 106 to obtain a data item. In the example shown in FIG. 3B, the data item, i.e. the decoded output, is a generated image.

[0084] FIG. 4 is a flow diagram 400 of an example process for generating a sequence of data elements. The process of FIG. 4 may be implemented by one or more computers in one or more locations; for convenience the process is described with reference to FIG. 1.

[0085] At step 402, the process obtains a current sequence of data element embeddings zi, Z2, ..., zn-i, that comprises a respective data element embedding of each data element at a position that precedes the current position, Each data element embedding represents a value of the data element that is a continuous variable.

[0086] At step 404, the current sequence of data element embeddings, zi, Z2,..., zn-i, is processed using a neural network comprising a sequence of attention neural network layers and an output block, to generate the data element znat the current position. The neural network can be the neural network 102 of FIG. 1.

[0087] At step 406, the output block processes the attention layer output of a last attention neural network layer before the output block to generate parameters defining a continuous distribution over possible values of a data element, e.g. parameters of a single or multivariate Gaussian distribution. For example a multivariate Gaussian distribution (GMM) might be used to model a complex distribution such as an image whereas a single Gaussian might be used for a dense prediction task such as panoptic segmentation.

[0088] At step 408, the output block samples from the continuous distribution to obtain the value of the data element at the current position, zn, which is a continuous variable. Any suitable sampling method may be used to sample from the neural network 102. In some implementations the sampling from the neural network 102 can be adapted, e.g. to modify a distribution of the generated data items.

[0089] As one example, the system can modify a variance of the continuous distribution by a scaling factor before sampling from the continuous distribution (“variance scaling”).For example the covariance matrices of the predicted Gaussian distributions can be scaled by a factor of greater than or less than unity to, in effect, change a temperature of the sampling process and adjust sample quality and diversity.

[0090] In some implementations the system can truncate the continuous distribution before sampling from the continuous distribution. More particularly the system can truncate the predicted distributions per dimension, thereby choosing a higher-density support. This can also be used to adjust sample quality and diversity.

[0091] In some implementations sampling from the continuous distribution to obtain the value of the data element at the current position can involve sampling a plurality of possible values for the data element at the current position. The system can thus implement a form of beam search. This can be done by maintaining a plurality of possible current sequences of data element embeddings whilst generating the sequence of data elements. Each possible current sequence of data element embeddings can comprise one or more of the possible values for a data element at one or more positions in the sequence.

[0092] This can done, for example, by retaining a subset of the possible current sequences of data element embeddings that have a greatest probability at a sampling step. After sampling from the continuous distribution for a final position in the sequence of positions one of the possible current sequences can be selected as the generated a sequence of data elements. For example B beams can be maintained, at every step sampling a number of candidates for every beam (a “fans”). The cumulative log probability for all beams and fans up to the current sampling step can be computed, and the B beams with the highest cumulative log probability can be selected.

[0093] As previously described, there are various ways in which conditioning data can be incorporated into the system. Apart from directly conditioning on a conditioning input, the conditioning input can be used in a modified version of the classifier free guidance used with diffusion models. This can involve effectively combining two continuous probability distributions, one with and one without conditioning data. This can improve the quality of generated data items.

[0094] More particularly processing the current sequence of data element embeddings using the neural network to generate the data element at the current position may involve i) processing the current sequence of data element embeddings with the neural network conditioned on a conditioning input that defines one or more characteristics of the generated sequence of data elements to generate parameters defining a firstcontinuous distribution over possible values of the data element, and ii) processing the current sequence of data element embeddings with the neural network conditioned on a null conditioning input (e.g. no conditioning input or a conditioning input of zero) to generate parameters defining a second continuous distribution over possible values of the data element. A value of a data element, e.g. at a current position in the sequence, can then be obtained by sampling from a combined continuous distribution defined by a combination of the first continuous distribution and of the second continuous distribution. In practice one way of doing this is by using rejection sampling.

[0095] For example, the system can sample from a combined probability densitywhere w is a weight hyperparameter and p(z|c) and p (z 10) are the continuous distributions with and without conditioning respectively. One way of sampling from PCFGZ\Cis to use rejection sampling, which involves sampling (in general multiple times) from a proposal distribution, e.g. a version of p(z|c), and then deciding whether or not to accept a sample based on PCFG .Z\C- The version of p(z|c) can be, e.g. a version with increased variance,2ac) where .c, 2acare the parameters predicted by the neural network 102 when given the conditioning input c.

[0096] FIG. 4 illustrates an example of autoregressive sampling from the neural network 102. One particular technique for generating a data item, e.g. an image, by sampling in parallel from the neural network 102 is now described. This follows the approach described in Chang et al. “MaskGIT: Masked Generative Image Transformer”, arXiv:2202.04200vl, Feb 2022.

[0097] This is an iterative approach in which, at each iteration, the neural network 102 predicts all the data elements, i.e. soft tokens, in parallel, but only keeps the most confident ones. The remaining data elements (soft tokens) are masked out and processed by the neural network 102 so that they are re-predicted at the next iteration. The process starts with all the data elements (soft tokens) masked out and can finish, e.g. when all data elements (soft tokens) are predicted or after a fixed number of iterations. A masking schedule can be applied to gradually reduce a fraction of masked data elements (tokens) at each iteration, e.g. a cosine schedule. It will be appreciated that such an approach can be particularly useful for inpainting or outpainting (extrapolation) of a data item such as an image or audio, e.g. where just the regions to be filled in are masked initially (and where a conditioning input can define the type of data item to be inpainted or outpainted).

[0098] In implementations of the described system the neural network 102 is trained to predict the probability of each data element (soft token), and to minimize the likelihood of the masked data elements (tokens), e.g. using a loss that is the sum of the negative log likelihoods of each of the masked data elements. One way of encoding the masked data elements (soft tokens) is as follows: The value of every masked data element (soft token) is set to zero, e.g. by replacing the locations in the data elements (e.g. z) corresponding to a mask M with zeros. Additionally one of two particular vectors can be concatenated in the feature dimension, a [MASK] vector for masked locations, and an [UNMASK] vector otherwise. The values of the particular vectors can, e.g. be learned.

[0099] FIG. 5 is a flow diagram 500 of an example process for generating image data, the image data specifying values for pixels of an image. The process may use the neural network 102 as described herein. The process of FIG. 5 may be implemented by one or more computers in one or more locations.

[0100] A sequence of data elements is generated 502 that comprise a respective continuous valued data element at each position in a sequence of positions. The sequence of data elements is generated, for each position after a first position in the sequence of positions, by steps 504 to 510.

[0101] At step 504, a current sequence of data element embeddings zi, zz, ..., zn-i is obtained that comprises a respective data element embedding of each data element at a position that precedes the current position. Each data element embedding represents a value of the data element that is a continuous variable.

[0102] At step 506, the current sequence of data element embeddings zi, Z2,. . ., zn-i is processed using a neural network, such as neural network 102, to generate the data element at the current position zn. The neural network comprises a sequence of selfattention neural network layers and an output block, and each self-attention layer has an attention layer input for each element of the sequence of data elements and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output.

[0103] At step 508, the output block processes the attention layer output of a last selfattention neural network layer before the output block to generate parameters defining a continuous distribution over possible values of a data element.

[0104] At step 510, the output block samples from the continuous distribution to obtain the value of the data element at the current position zn. The value of the data element at the current position znis a continuous variable.

[0105] At step 512, the sequence of data elements z is processed using a trained decoder neural network, such as decoder 106, to generate image data defining values of pixels of an image.

[0106] FIG. 6 is a flow diagram 600 of one example process for training a neural network, such as neural network 102, for generating a sequence of data elements. The process of FIG. 6 may be implemented by one or more computers in one or more locations.

[0107] At step 602, in this example process, an autoencoder, e.g. a trained autoencoder, is obtained. The autoencoder comprises an encoder neural network and a decoder neural network. An example of the encoder neural network and a decoder neural network are the encoder neural network 104 and decoder neural network 106 respectively. The trained autoencoder can be obtained by pre-training the autoencoder and then using the pre-trained autoencoder in the process; or the trained autoencoder can be obtained by training the autoencoder jointly with the neural network 102. Joint training can involve backpropagating gradients of a training objective function through the decoder neural network 106, the neural network 102, and the encoder neural network 104.

[0108] At step 604 the neural network is trained by, for a plurality of the training data items, processing 606 the training data item using the encoder neural network to generate a first set of parameters defining the posterior distribution of the set of latent variables z.

[0109] The posterior distribution is sampled 608 from to determine training sampled values of the set of latent variables z, and the neural network is trained 610 using the training sampled values of the set of latent variables z. That is, in this example the neural network 102 is trained to reconstruct the values of the set of latent variables z from the encoder neural network 104, optionally incorporating conditioning data as previously described.

[0110] There are many different types of objective function that may be used when training the neural network. As one example the model may be trained using a softmax cross entropy loss, e.g. using language model style teacher forcing with a softmax cross entropy loss as previously described. As another example the neural network may betrained with a masking loss, e.g. a loss that requires the model to predict masked-out data elements (soft tokens).

[0111] The neural network 102 can be trained to facilitate using classifier free guidance as described above. This can involves randomly discarding the conditioning data. More specifically, when training the neural network 102 using conditioning data this conditioning data is randomly discarded, e.g. by using a null class, e.g. setting c = 0, so that the neural network 102 is trained using both class-conditional and unconditional training data items.

[0112] As a particular example of training the neural network 102, consider the previously described ^-mixture Gaussian mixture model (GMM) for which the neural network 102 predicts 2kd + k parameters per soft token (kd mean and kd variance parameters and k mixture probabilities). For a batch of B training data items represented by the encoder neural network 104 as -dimensional features sequences of length L (e.g. 1 < d < 64), the neural network 102 output is a B x L x (2dk + fc) tensor y = [m; s; n] where the mean m and standard deviation s are B x L x dk tensors and n is a B x L x k tensor; all stacked along the last dimension, e.g. m composed as [m1; ... ; mk] where the mlare B x L x d mean tensors stacked along the last axis; 5 lower-bounded to a small positive constant; and the mixture weights are normalized, e.g. with a softmax so that they sum to 1, i.e.b> I. Then the loss to be minimized is given by the negative log likelihood --logp (z&), which can be expanded aswhere Jf(z|m, s) is a Gaussian density with standard deviation s (variance s2), i labels mixture components, and j labels output feature channels. Heresbllare predicted from zb l, ... , zb t-1and zb Qcan be set to a learned [CLS] vector. This can be straightforwardly evaluated, e.g. using the JAX library distrax as:# With z and m, s , pi as above pdf = distrax . MixtureSameFamily ( mixture distribution=distrax . Categorical (logits=pi) , components distribution=distrax . MultivariateNormalDiag ( loc=m, scale diag=s))# Sample from next token distribution (teacher forcing) samples = pdf . sample ()# Compute NLLloss = -pdf. logjprob (z) . mean ( )

[0113] When using a single Gaussian instead of a GMM this becomes:

[0114] In both of the above examples, where the sequence of data elements is processed in parallel, and the training involves masking data elements (soft tokens) as previously described, the sum over L is reduced to a sum over indices I of the masked locations, and the loss can be normalized by the number of masked elements.

[0115] As previously mentioned, in general training comprises backpropagating gradients of an objective function, e.g. a loss function, to update learnable parameters, e.g. weights, of the neural network. This may use any appropriate gradient descent optimization algorithm, e.g. Adam or another optimization algorithm.

[0116] The training data may be any data appropriate to the task being performed. There are many publically available datasets for the types of task listed below. The training data may be representative of real world data. For example, when the training data comprises images, the images may be of a real world environment, captured by a physical imaging device (e.g. camera). When the training data comprises audio, the audio may be of a real world environment, captured by a sound capture device (e.g. microphone).

[0117] To recap, the decoder neural network and the encoder neural network may be trained on training data items of the same type as are to be generated, optionally on also on conditioning data. For example the training data items may comprise labelled training data items where the conditioning data, e.g. label for a data item, specifies, e.g. a class or category of data item to be generated. The decoder neural network may have been trained by training an autoencoder comprising the decoder neural network, e.g. VAE or ^-VAE.

[0118] Training such an autoencoder may comprise, for a plurality of training data items, processing the training data item using a encoder neural network to generate a first set of parameters defining a posterior distribution of the set of latent variables, sampling from the posterior distribution to determine sampled values of the set of latent variables, and processing the sampled values of the set of latent variables using the decoder neural network to generate an output training data item. The encoder neural network and the decoder neural network can be trained jointly using an objectivefunction that has a first term dependent upon a difference between the training data item and the output training data item and a second term dependent upon a difference between the posterior distribution and a prior distribution for the set of latent variables.

[0119] Each of at least some of the training data items may include conditioning data for the training data item that defines one or more characteristics of the respective training data item. Training the neural network 102 using the training sampled values of the set of latent variables can involve training the neural network whilst conditioned on a conditioning input comprising the conditioning data for the training data item used to generate training sampled values of the set of latent variables. The processing of the set of latent variables using the decoder neural network to generate the data item can also be conditioned on a conditioning input comprising the conditioning data.

[0120] In some implementations the training involves obtaining mask data, the mask data specifying locations to be masked in the training sampled values of the set of latent variables; and replacing, based on the mask data, values stored in the specified locations in the training sampled values with zero to obtain masked training sampled values. The neural network can be trained using the masked training sampled values of the set of latent variables. This can also involve concatenating a data element embedding corresponding to a latent variable and a vector with the masked training sampled values of the set of latent variables. The vector can indicate a location of masked values and / or a location of unmasked values.

[0121] In some implementations the system 100 generates a data item by generating a set of latent variables and processing the set of latent variables using a decoder neural network to generate the data item. Generating the set of latent variables can involve generating a sequence of data elements as previously described, each data element defining a value of a respective latent variable.

[0122] In some implementations a small invertible (normalizing) flow model, also referred to herein as an adapter or adaptor block, can be used to better match the latent distributions of the encoder neural network 104 and the distribution predicted by the neural network 102. For example such an adapter can be used to map the VAE latent sequences to a new latent space of the same dimensions. Such an adapter may comprise a “volume preserving” additive coupling layer-based model and may have a diagonal Jacobian. The neural network 102 is then trained jointly with the adapter to predict the sequences in this transformed latent space induced by the adapter (using the same loss; i.e. no additional log-determinant term is needed). At inference time, samples drawnfrom neural network 102 are first processed by the inverted adapter, and then decoded with the VAE decoder 106. The adapter does not require additional losses thanks to invertibility and adds a negligible compute and model parameter overhead.

[0123] An example of an adapter (block), i.e. invertible normalizing flow model, is described in greater detail below. The example is described as processing the set of latent variables; the same approach can be used for the normalizing flow model used later to process the data element embeddings.

[0124] Thus processing the set of latent variables using the decoder neural network to generate the data item can involve processing the set of latent variables to determine a modified set of latent variables then processing the modified set of latent variables using the decoder neural network. The modified set of latent variables can be determined by an (inverse) adapter block, e.g. one that performs operations as described below.

[0125] In implementations each latent variable comprises a vector and processing the set of latent variables comprises partitioning each vector into a first part (x-^ = x1;d) and a second part (x2= xd;D). A first part of a corresponding modified vector can be determined from the first part, e.g. it may be the same as the first part. A second part of the corresponding modified vector can be determined by modifying the first part using a learned transformation to obtain a modified first part. The modified first part and the second part may then be combined, e.g. by subtraction. Here the vector and the modified latent vector may each be D-dimensional.

[0126] The adapter block (normalizing flow) and inverse adapter block (inverse normalizing flow) may be implemented using one or more neural network layers. For example in some implementations modifying the first part using a learned transformation comprises processing the first part using a (trained) adapter neural network to generate the modified first part. The adapter neural network may have the same input shape and output shape, e.g. w x d x h (e.g. in an implementation as described above the adapter can be applied to the VAE latent z before re-shaping it into a sequence).

[0127] The adapter (normalizing flow) neural network may have any suitable architecture e.g. one or more feedforward, convolutional, or attention based neural network layers. As a particular example the adapter neural network may comprise one or more convolutional, bijective blocks where each block has one or more convolutional layers, optionally interleaved with normalization layers. For example a stack of one or more (as an example, 8) convolutional bijective i-RevNet blocks can be used, e.g. withhidden channel dimension 4d (Jacobsen et al. “i-RevNet: Deep invertible networks”, ICLR 2018).

[0128] In some implementations modifying the first part using a learned transformation comprises determining a scaling and an offset based on the first part, and applying the scaling and offset to the second part (x2). For example the scaling may comprise a scaling determined dependent on the first part (s^)) and an offset dependent on the first part (t(xx)).

[0129] The inverse adapter block can provide an inverse operation to an adapter block used during training (as described later). For example an adapter block may provide an outputany invertible map e.g. g(a; b) = a + b and rn x- ) is a function, e.g. an affine transform or a function defined by the adapter neural network. The inverse adapter block may provide an output

[0130] Such approaches can improve processing, e.g. by making it easier for the neural network comprising the sequence of attention neural network layers to model a set of latent variables that can be decoded using the decoder neural network, in particular facilitating the representation of more complex distributions. The particular approach used for determining the modified set of latent variables is invertible, and thus the same adapter neural network can be used both an adapter block during training and the inverse adapter block during inference.

[0131] For example, in cases where the neural network is trained to match the latent variables of another neural network, such as an encoder of an autoencoder, such as VAE, it can be beneficial to use an adapter to map the latent variables of the encoder into a new latent space, and train the neural network on the modified latent variables in the new latent space. An inverse adapter can be used to map the latent variables back into the original latent space.

[0132] In training, as an example, sampling from a posterior distribution as described above to determine training sampled values of the set of latent variables comprises sampling from the posterior distribution to determine an intermediate training sampled values of the set of latent variables, and processing the intermediate training sampled values of the set of latent variables to determine the training sampled values of the set of latent variables. In implementations each training sampled value of the set of latent variables comprises a vector. Processing the intermediate training sampled values ofthe set of latent variables can involve partitioning each vector into a first part and a second part, and can also involve determining a first part of a corresponding modified vector from the first part. A second part of the corresponding modified vector can be determined by modifying the first part using a learned transformation to obtain a modified first part, and combining the modified first part with the second part.

[0133] Processing the intermediate training sampled values of the set of latent variables to determine the training sampled values of the set of latent variables may comprise processing the training sampled values using an adapter block, e.g. as described above, e.g. comprising the above described adapter neural network.

[0134] In this way, the training sampled values of the set of latent variables may be mapped to a new latent space of the same dimensions. Mapping to different latent spaces can improve processing. For example, the neural network may be able to learn the new latent space more efficiently than the original latent space of the encoder. At inference, an inverse adapter, as described above, may be used to map the latent variables back into the original latent space such that the decoder can process them.

[0135] The first part of a corresponding modified vector and the second part of a corresponding modified vector may be combined to provide a modified latent vector. The combining may be performed using an inverse operation to that described above. For example, combining the modified first part with the second part may comprise an addition operation or a multiplication operation. As previously described, in implementations the vector and the modified latent vector each comprise a D- dimensional vector.

[0136] Modifying the first part using a learned transformation may comprise processing the first part using the trained adapter neural network to generate the modified first part, e.g. by determining a scaling and an offset based on the first part, the scaling and offset comprising the modified first part, and applying the scaling and offset to the second part.

[0137] The system 100 can be used for performing a dense prediction task. In a dense prediction task. A dense prediction task generally involves assigning a label or value to each element of a data item, e.g. to each pixel of an image or to each sample of an audio waveform, e.g. to label the pixel / sample with depth for a depth prediction task, an object identifier for a segmentation task such as semantic or panoptic segmentation, a pixel color for a colorization task, and so forth. Thus, for example, a data item generated by the system 100 may comprise image data, the image data specifyingvalues for pixels of an image. The values for the pixels may define the result of a dense prediction task.

[0138] The system 100 can be configured to perform a dense prediction task by using the encoder neural network 104 to generate an input for the neural network 102. The encoder neural network 104 and decoder neural network 106 can have been trained in combination as an autoencoder, e.g. a ( L)VAE, to generate the dense prediction output, e.g. to process an image to generate a segmented image output, a depth prediction output, a color output, and so forth.

[0139] The neural network 102 is configured as an encoder-decoder Transformer. This can comprise an encoder Transformer to receive and encode a sequence of data element embeddings that represents a data item to be processed, and a decoder Transformer to process the encoded a sequence of data element embeddings and to use this as conditioning data when autoregressively generating the sequence of data elements that represents the dense prediction output. The system can be trained as previously described.

[0140] As a particular example, such an approach can follow the UViM framework described in Kolesnikov et al. “UViM: A Unified Modeling Approach for Vision with Learned Guiding Codes”, arXiv:2205.10337v3, Oct 2022), but using the system described herein rather than a conventional discrete vocabulary transformer and vector quantized VAE. For example, where the dense prediction task is performed on an image, e.g. an RGB image, the encoder neural network 104 may comprise a pre-trained Vision Transformer (Dosovitskiy et al., arXiv:2010.11929). The encoder part of the neural network 102 can omit causal masking; the decoder part of the neural network 102 can use causal masking and can insert a cross-attention layer after each selfattention layer to attend to the encoder part of the neural network.

[0141] Although some implementations of the system are based around use of an autoencoder, this is not essential. In principle a system as described above could model an image directly, e.g. represented as pixels or as flattened image patches. In practice modelling in pixel space would require many soft tokens; and modelling flattened image patches would involve modelling a complex distribution, which is difficult and would need a large number of Gaussian mixture components.

[0142] To ameliorate this the system can include a normalizing flow model, e.g. implemented in the same way as the previously described adapter block. As previously described, this is a model that is invertible (bijective), and that converts one probabilitydistribution into a second different probability distribution. Here flow refers to the path traversed by random variables (and refers to a characteristic that the invertible transformation can be composed). The flow is normalizing because it results in a valid probability distribution (the model output has a normalized (probability) density. In practice such a model can be implemented using one or a few invertible coupling blocks, i.e. as previously described, (with the same input and output dimension). In general the normalizing flow model is lossless, and differentiable.

[0143] As previously mentioned there are many types of normalizing flow models, e.g. planar flow models, radial flow models, autoregressive flow models, models based on Langevin or Hamilton flow, and so forth (see, e.g., the previous references, Rezende et al., “Variational Inference with Normalizing Flows”, arXiv: 1505.05770, 2016; Ho et al., “Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design”, ICML 2019; and others). The normalizing flow model generally corresponds to the adapter block as described above, but can be a larger model, e.g. using a neural network with more layers to implement the learned transformation.

[0144] Such an invertible flow model can, e.g., convert a sequence of flattened image patches into a sequence of latent vectors with the same dimension. Thus an invertible normalizing flow model can be used as the encoder neural network 104, and the inverse normalizing flow model, i.e. the same normalizing flow model inverted (swapping the input and output), can be used as the decoder neural network 106. The normalizing flow model in effect learns a soft token representation of the input, e.g. image. In implementations such a normalizing flow model (in effect the encoder neural network 104 and the decoder neural network 106) can be trained end-to-end with the neural network 102.

[0145] This can be extended to modelling multimodal data, e.g. interleaved images and text. Text tokens (which can be embedded in a discrete space) can be predicted via a separate prediction head (including a softmax) on, e.g., the same neural network 102, and images can be delimited, e.g. with [BOI] (beginning of image) and [EOI] (end of image) tokens to indicate when to use which prediction head. For text-to-image generation, the model can be prefilled with a text prompt and [BOI] to then sample image latents. Conversely, vision-language understanding tasks such as captioning and VQA (Visual Question Answering) can be tackled by prefilling image latents obtained by applying the flow model, and optionally a text prompt, and then sampling text tokens(the model having been trained on appropriate interleaved data). Since the invertible flow model preserves all the input information, such a multimodal model can in principle learn any task. As an example, visual text (text in an image) can be modelled whilst preserving the text information, facilitating tasks such as OCR (optical character recognition).

[0146] Thus in some implementations a sequence of data element embeddings may be processed using a normalizing flow model to generate a corresponding sequence of processed elements at respective positions to the current sequence of data element embeddings, each processed element comprising a vector with a dimensionality matching the dimensionality of the corresponding data element. Following this, the sequence of processed elements may be processed with a decoder neural network to generate parameters defining a continuous distribution over possible values of a data element.

[0147] The sequence of data elements may be generated following the receipt of a prompt sequence comprising either one or more tokens representing pixel values of an image, or one or more tokens representing a sequence of text. The generated sequence of data elements may then represent either, respectively, a sequence of text that characterizes content of the image (such as, for example, a caption, or an answer to a question about or expressed in the image), or pixel values of an image characterized by the sequence of text.

[0148] The system 100 comprising the normalizing flow model, e.g. comprising the encoder neural network 104 and decoder neural network 106 implemented using the (same) normalizing flow model, can be trained broadly as previously described, but with the addition of a log-determinant term that arises from the normalizing flow model. The log-determinant term can be the log of the absolute value of the determinant of the Jacobian of the transformation applied by the normalizing flow model.

[0149] For example the system can be trained using a sequence of continuous-valued data elements by processing each data element using a normalizing flow model to generate a respective training data item, thereby obtaining a sequence of training data items. Each training data item can comprise a vector with a dimensionality matching a dimensionality of the respective continuous-valued data element. The neural network can be trained to maximise a likelihood of a next element in the sequence of training data items. The continuous-value data elements may, e.g., represent pixels of an image. Merely to illustrate, an H x W x 3 image, represented as an HW / p2x 3p2sequence xof flattened p x p image patches, can be transformed by the normalizing flow model into a latent sequence of the same shape.

[0150] As an example, consider modelling a data item x, such as an image (modelling pixel space). The normalizing flow model, f (x) losslessly maps x to a sequence of embeddings, f (x) = {z1(zn), i.e. to a sequence of soft tokens, preserving the total number of input dimensions. These embeddings are modelled autoregressively by the neural network 102, as p(z) = I ?=i P(zdz<t) where z<t= Zj_1?..., z (or, in a conditional model, determining p(z|c)), using a Gaussian mixture model. A training objective for training the system end-to-end can then be to maximize the log-likelihood lower bound z xL(x) = logp( (x)) + log detTraining the system end-to-end can refer to training both the neural network 102 and the normalizing flow model using the training objective, e.g. by backpropagating gradients the training objective through the inverted normalizing flow model, the neural network, and the normalizing flow model. Optionally uniform “dequantization” noise can be added to the data items, e.g. as x = / + u where u~[0,l]. During inference the neural network 102 generates a sequence of soft tokens that are decoded to a data item using the normalizing flow model with an inverse of the flow.

[0151] In some implementations only a proper subset (i.e. less than all) of the output dimensions of the invertible flow model is processed (autoregressively) by the neural network 102. In particular relatively lower frequency information can be processed and relatively higher frequency information can be modelled with a Gaussian (standard normal) distribution, p^. That is the embeddings from the normalizing flow model can be divided into low frequency components, z, and high frequency components, z, f (x)=I X • The training obj ective can then be to maximize: L(x) = logp(z) + p^(z) + log det

[0152] In some implementations a data item, e.g. an image, can be re-shaped into a sequence of flattened patches and then a learnable, invertible linear map W can be applied along the channel dimension, to learn to separate important dimensions of the flattened patches from substantially redundant ones. This can be done by feeding the first d channels of the map output xWTto the normalizing flow model and modeling the remaining channels with a Gaussian distribution. Then optimizing the trainingobjective, e.g. maximizing the above objective or minimizing the negative log likelihood (NLL), will tend to ensure that the hardest to model part of the sequence is modelled by the neural network 102, whereas the low-level noise will be mapped to the Gaussian prior.

[0153] In some implementations classifier free guidance can be used during training, as previously described.

[0154] In some implementations the training process can include a “noise curriculum”, adding Gaussian noise to elements, e.g. pixels, of a data item during the training, where the added noise is greatest at the start of the training decays towards zero, e.g. according to a cosine decay schedule for the noise standard deviation. This can facilitate the neural network 102 gradually learning to model finer levels of detail whilst retaining previously learned patterns. This noise curriculum can be considered as a form of data augmentation.

[0155] Thus the training can involve generating noise values from a Gaussian distribution, augmenting each training data item with a respective noise value, and over the course of training the neural network, modulating a magnitude of the Gaussian distribution according to a decreasing function, e.g. a cosine function. The use of noise modelled by a decreasing function has the effect that noise is more prominent at the beginning of the training process, and becomes less prominent during training.Towards the beginning of training, the neural network may only be able to resolve and leam coarse information in a data item, e.g. image; reducing the magnitude of the noise then trains the network to handle progressively finer details. This technique can be used generally for training the neural network 102, in the context of any of the methods described herein.

[0156] Merely as an example, for an integer-valued RGB image / , a noisy image can be obtained as [ / + crtlV(O,l)], where the noise scale atas a function of the training1 " cos t / r progress t E [0,1] can follows the cosine schedule at= a0— ~ — •

[0157] FIG. 7 illustrates a process of training a model comprising the neural network 102 and an invertible model 700, e.g. a normalizing flow model as described above, in the example using teacher forcing. In the illustrated example the ground truth comprises text tokens and images. The images are converted into soft tokens using the invertible (normalizing flow) model 700. The sequence of tokens is processed autoregressively by the neural network 102, which generates outputs that parameterizeeither a discrete distribution for generating discrete, text tokens, or a Gaussian mixture for generating soft tokens.

[0158] In the illustrated example a prompt (“<BOS> Jet Former <BOI>”) is provided, the invertible flow model 700 is used to encode the image into soft tokens, and the neural network 102 is trained to autoregressively predict the soft token distribution. In another example (not shown), the image precedes the text, and the neural network 102 is trained to autoregressively predict the token tokens (using teacher forcing and a NLL loss), to predict a caption for the image. The process can also be used to train the system to predict, e.g., segmentation maps or depth maps.

[0159] FIG. 8 illustrates an inference process in which given a text prompt as conditioning data the (trained) system is used to autoregressively sample soft tokens from the output of the neural network 102, that are decoded into an image using the invertible (normalizing flow) model 700.

[0160] FIG. 9 is a flow diagram of a further example process for generating a sequence of continuous-valued data elements using a trained neural network, such as neural network 102, in combination with a trained invertible (normalizing flow) model. The process of FIG. 9 may be implemented by one or more computers in one or more locations. The continuous-valued data elements can be values for pixels of an image, e.g. pixel color or intensity values.

[0161] The process autoregressively extends a sequence of soft tokens, each soft token comprising a vector of continuous values (step 900). Each soft token can comprise a continuous value or a vector of continuous values that represent a patch of a sequence of patches of the image, e.g. that tile the image, or a region of some other type of data item.

[0162] A current sequence of soft tokens is processed using a neural network to generate a set of parameters of a continuous distribution, e.g. parameters of a multivariate Gaussian Mixture Model as previously described (step 902). The neural network can comprise a sequence of attention neural network layers and an output block, e.g. as previously described for neural network 102; it may comprise a Transformer neural network. Implementations of the technique process each soft token using one more embedding layers, e.g. linear embedding layers, to generate a data item embedding, and the data item embeddings are processed using the Transformer neural network.

[0163] The process can sample from the continuous distribution to obtain a next soft token to extend the current sequence of soft tokens (step 904), e.g. until a sequence representing a complete data item, e.g. image, is obtained.

[0164] Each soft token of the sequence of soft tokens can be processed using a trained invertible model (step 906), which acts as a decoder, to generate a set of values for elements of the generated data item, e.g. pixel values for each respective patch of the image.

[0165] The invertible model may comprise one or more invertible coupling blocks. In implementations the invertible model has been trained to process an, e.g. a d- dimensional, input vector using a learned invertible transformation to generate an output vector with the same number of dimensions, e.g. also -dimensional. In general the (trained) invertible model is a model that applies an invertible transformation (mapping) between the input vector and the output vector, e.g. using a learned mapping that is parameterized by one or more learned parameters. One class of such a model is known as a normalizing flow model.

[0166] In implementations the invertible flow model converts a sequence of flattened image patches into a sequence of latent vectors with the same dimension. The Transformer neural network is then trained to predict this latent sequence. At inference time, a sequence of latents is sampled autoregressively from the Transformer neural network, and decoded into image pixels by applying the inverse flow.

[0167] Implementations of the system can, without modification, process tokens from a discrete vocabulary as well as soft tokens. As an example text tokens can be generated using a SentencePiece tokenizer.

[0168] Some implementations of the method involves processing a prompt sequence. The prompt sequence can, e.g., comprise text which may be represented by tokens selected from a discrete vocabulary of tokens. Then the image can be an image that is characterized by the text, or that performs a dense prediction task instructed by the text. Here a dense prediction task can be one that assigns a value to each pixel, e.g. a depth value (for a depth prediction task), or a label denoting an object class or instance to which the pixel belongs (for a semantic or panoptic prediction task).

[0169] The system can be prompted e.g. with text, or audio, or data representing actions of a mechanical agent such as a robot. This can involve receiving an initial sequence of tokens representing data items of a different modality to images, e.g. text, andprepending the initial sequence of tokens to an initial version of the current sequence of soft tokens.

[0170] A beginning of image (BOI) token can be added at the end of the initial sequence of tokens to signal the start of image generation to the system, to facilitate generation of soft tokens representing an image. Where token generation continues after image generation, or where a prompt includes an image, an end of image (EOI) token may be included to signal the end of an image to the system, e.g. to thereafter facilitate generation of tokens from a discrete vocabulary representing text. A beginning of sentence (BOS) token can be added at the start of a sequence of text tokens to signal the start of text generation to the system.

[0171] Some implementations of the system include multiple token prediction heads for generating output tokens. For example a first, image generation token prediction head, comprising the trained invertible flow model, can be used for image generation. A second token prediction head can be used for processing tokens representing data items of a second, different modality to images, e.g. to generate text, e.g. by providing an output comprising a set of scores representing a probability of each token in a vocabulary of tokens, from which a token to continue a current sequence is selected.The process can then involve detecting a mode identification token (e.g. a BOI or EOI token as previously described), e.g. in the current sequence or in the initial sequence. In response the system can select a suitable token prediction heads for generating a next (output) token.

[0172] In some implementations processing each soft token using the trained invertible flow model comprises reducing a dimension of the vector of continuous values representing a soft token prior to processing each soft token of the sequence of soft tokens using the trained invertible flow model. This can involve selecting a subset of the dimensions, optionally after processing each soft token using a (leamed / leamable) linear, invertible map. The remaining dimensions can be modelled as a predetermined (prior) distribution such as a Gaussian distribution. This can facilitate the model focussing computation on more significant, low-level or semantic information, e.g. to reduce overall computational requirements.

[0173] FIG. 10 shows examples of data items generated by the systems described above. FIG 10A shows examples of class conditional images generated by the system of FIG. 1. FIG. 10B shows examples of a panoptic image segmentation task performed by a system as described above. FIG. 10C shows an example of a system that uses anormalizing flow model similar to FIGS. 8 and 9, generating an image in response to a text description.

[0174] A task performed by the neural network 102 can be to generate a sequence of data elements as described above, either unconditionally, e.g. in accordance with a training distribution, or conditionally, as described further below.

[0175] The sequence of data elements representing an entity i.e. a data item, can include a respective data element at each position in the sequence of data elements. Generally, the sequence of data elements can represent any appropriate entity, e.g., any appropriate type of data, and can include any appropriate number of data elements, e.g., 1 data element, 10 data elements, 100 data elements, 1,000 data elements, 10,000 data elements, 1 million data elements, 5 million data elements, or any other appropriate number of data elements.

[0176] In one example, each data element can represent a pixel in an image, e.g. an intensity value of the pixel, and the sequence of data elements can collectively represent the image.

[0177] As another example, each data element can represent an audio sample in an audio waveform, e.g. a time or time-frequency sample, and the sequence of data elements can collectively represent the audio waveform. As another example, each data element can represent a musical note, and the sequence of data elements can collectively represent a musical composition.

[0178] As another example, each data element can represent a monochrome, color, or hyperspectral pixel in a respective still or moving image, e.g. a video frame of a video, e.g. an intensity value of the pixel. The sequence of data elements can collectively represent the image, e.g. video.

[0179] As used herein an “image” includes a point cloud e.g. from a LIDAR system that captures an image of the real world, and a “pixel” includes a point of the point cloud. Similarly “video” includes a time sequence of point clouds. For example, each data element can represent the location of a point in 3D space, e.g. in x, y, z coordinates, and the sequence of data elements can collectively represent a point cloud. The point cloud can characterize a 3D geometry of an environment, e.g., where each point in the point cloud represents a point on a surface or object in the environment.

[0180] As another example, each data element can represent a respective structure parameter from a plurality of structure parameters that collectively define a structure of a protein. For example for each position in an amino acid sequence of the protein acorresponding data element may define a set of structure parameters that characterize a three-dimensional spatial configuration of the amino acid at the position. The set of structure parameters can comprise, e.g., torsion angles of the bonds between the amino acids in the protein, e.g. backbone torsion angles, or three-dimensional (3D) coordinates defining the positions of one or more atoms in the amino acid, e.g. nitrogen, alpha carbon, and beta carbon atoms of the protein.

[0181] As another example, each data element can represent a continuous action, e.g., that can be performed by an agent interacting with an environment. For example data elements can specify positions, torques, or other control signals for the parts of a mechanical agent, or higher-level control commands. The sequence of data elements can collectively represent a sequence of actions that can be performed by an agent to interact with an environment. The agent can be, e.g., a mechanical agent, e.g., a robot or a self-driving vehicle, and the environment can be a real-world environment. The actions may selected in response to one or more observations of the real-world environment, e.g. to control the mechanical agent to perform a particular task such as moving, or manipulating an object.

[0182] As another example the input sequence can represent a time series and the output sequence may comprise a continuation of the time series. For example the input sequence may be a sequence representing the output of an electricity generating plant, e.g. a solar or wind electricity generating plant, or a sequence representing electricity consumption, and the output sequence may provide a forecast of the electricity generated or consumed. As another example the input sequence may be a sequence representing a level of traffic on one or more roads and the output sequence may provide a forecast of the future traffic.

[0183] In implementations, the neural network can be conditioned on a conditioning input comprising data that specifies one or more desired characteristics of sequence of data elements, i.e. data item, to be generated by the neural network. As some examples, the neural network can be configured to process a conditioning input as well as a sequence of data element embeddings; as another example the conditioning input may comprise a prompt prepended to sequence of data element embeddings to be processed.

[0184] When the neural network 102 is not conditioned on a conditioning input it can be used to obtain an example of a sequence of data elements from the training distribution i.e. from the distribution of a set of training examples used to train the neural network, e.g. another image or sound like those the system has been trained on,or a protein or DNA sequence with properties similar to others that the system has been trained on. Thus it is not necessary for the neural network to be conditioned on data that specifies desired characteristics of sequence.

[0185] When the neural network is conditioned on a conditioning input the conditioning data can, e.g., characterize a sequence of text, and when conditioned on the conditioning data, the neural network can generate a sequence of data elements that represents a verbalization of the sequence of text, e.g. where each data element represents an audio sample in an audio waveform. Thus the system can perform a text-to-speech conversion task. Also or instead the conditioning data can identify a desired speaker for the audio, i.e., so that the system generates audio data that represents speech by the desired speaker. As another example the conditioning data can specify a classification for the audio waveform into a class from a set of possible classes, so that the system generates audio data that belongs to the class, e.g. a musical genre or instrument.

[0186] As another example, the conditioning data can define a set of properties of a protein (e.g., stability, solubility, etc.), and when conditioned on the conditioning data, the neural network can generate data defining a protein that is predicted to have the properties specified by the conditioning data. As some other examples, the conditioning data can specify a gene to be activated by a DNA sequence, or a regulatory property of the DNA sequence or protein, or a target binding site for the DNA sequence or protein. The data defining the protein can be, e.g. a protein structure or an amino acid sequence as described above. Such as system can be trained from real-world experimental data. The DNA sequence, e.g. protein, may then be physically synthesized.

[0187] As another example, the conditioning data can specify one or more features of an image (e.g., an object to be shown in the image and optionally its location), and when conditioned on the conditioning data, the neural network can generate an image having the features specified by the conditioning data. Thus the conditioning data can specify a classification for the image or part of the image into a class from a set of possible classes, so that the system generates an image or image part that belongs to the class. As another example the conditioning data specify the location of an object to be included in a generated image (which need not involve specifying the object). As another example the conditioning data can be a sequence of text and the output data item can be an image that describes the text, i.e., the conditioning input can be a caption for the output image.

[0188] As another example, the conditioning data can specify one or more future times or conditions and the he neural network can generate a predicted image representing thefuture times(s) or condition(s). This may then be used by a control system, e.g. using model predictive control, to control a mechanical agent such as a robot to perform a particular task, by processing the predicted image using the control system to generate control signal to control the mechanical agent, in accordance with the predicted image to perform the task. The mechanical agent may be a real world agent, e.g. an agent that operates in a real world environment.

[0189] As another example, the conditioning data can specify one or more features of a point cloud (e.g., an object characterized by the point cloud), and when conditioned on the conditioning data, the neural network can generate a point cloud having the features specified by the conditioning data. Examples of conditioning data used for point cloud generation are as those described above for image generation.

[0190] As another example, the conditioning data can specify text or one or more features of a sequence of text (e.g., a topic of the sequence of text), and when conditioned on the conditioning data, the neural network can generate a sequence of audio data values representing an audio waveform specified by the conditioning data, e.g. that is a spoken version of the text.

[0191] The system 100, e.g. the decoder neural network 106, can generate any type of data item; some examples are below. The neural network can be used to generate the set of latent variables either unconditionally, e.g. in accordance with a training distribution, or conditionally, as described elsewhere herein.

[0192] As one example the data item may comprise a 2D or 3D still or moving image, and the decoder may generate a value for each pixel in the image, e.g. an intensity value of the pixel. The image generation may be conditioned on a conditioning input that specifies the type of image to be generated, e.g. a type of object or action depicted in the still or moving image, and / or a location or viewpoint for which the image should be generated. For example the conditioning input can specify an object class from a plurality of object classes to which an object depicted in the image data item should belong.

[0193] In some implementations an image is processed, e.g. to identify one or more objects in the image, or to classify the image into one or a plurality of classes, or to generate a text description or caption for the image, or to answer a question about the image. In this case the image (or video) may have been captured using a camera i.e. captured from the real world. Objects in an image (or video) may comprise objects, e.g. physical objects, represented by the image (or video).

[0194] As another example, as previously described the system may be used to perform a dense prediction task on an input image and the generated data item may be an image comprising pixels (that may but need not correspond to pixels of the input image), with pixel values that are the output of the dense prediction task. The input image may be captured by a real world camera and may comprise an observation of the real world.

[0195] Some examples of dense prediction image processing tasks are: image segmentation, e.g. semantic segmentation or instance segmentation; depth prediction; image colorization; keypoint prediction; pose estimation, e.g. 3D pose estimation; surface normal estimation, e.g. by determining a vector in 2D or 3D; or object detection, including object tracking. Many other types of task may be performed in the same way, e.g. a curvature or other shape estimation task, a task that involves identifying aspects of an image using color, and so forth.

[0196] As one particular example, in an image segmentation task the pixel values for the task may each have a categorical value defining a category for the pixel, or a value representing a probability that the pixel belongs to a particular category. The category may represent an object or type of object or (for video) an action or type of action. For example in a semantic segmentation task the pixel values may identify a type or category of object and in an instance segmentation task the pixel values may (also) distinguish between different instances of the same category of object. Thus the set of pixel values for the input image pixels can locate categories or instances of objects or actions in an image. More generally a pixel value can distinguish between an object (or action) and image background, and the set of pixel values for the input image can, e.g., perform an object localization, detection, or tracking task, e.g. for gesture recognition. Merely as some illustrative examples, object segmentation may be used to segment medical images, to label pixels of an input medical image in accordance with whether they show a region of a human or animal body in which a particular medical condition is present.

[0197] As another example, in a (monocular) depth prediction task the pixel values may each comprise a scalar value representing an estimated depth value for the pixel, e.g. a distance of the pixel (for an object) from in a depth or z-direction from an x-y image plane or camera viewpoint. Or the pixel values may each define a depth distribution, e.g. a probability distribution over discrete depth value buckets. The generated pixel values can define a depth map for the input image.

[0198] As another example, in a keypoint prediction task the pixel values for the input image pixels may identify keypoints in the input image, e.g. by labelling a pixel as akeypoint or as one of multiple keypoints. The set of pixel values for the input image pixels can thus label keypoints in the input image, e.g. one or more keypoints of an object in the image. For example keypoints may define landmarks of an object represented in the image, e.g. the positions of body joints for a human.

[0199] As another example, in a pose estimation task the pixel values for the input image pixels may map the input image pixels to a 3D surface, e.g. of a human body or face. Or the pixel values may estimate a 6D pose representing translation and orientation components of an object in the input image, e.g. in quaternion form. The set of pixel values for the input image pixels can estimate the pose of one or more objects in the input image.

[0200] As another example, in a surface normal estimation task the pixel values for the input image pixels may comprise a vector in, e.g., three dimensions defining the surface normal. The set of pixel values for the input image pixels can provide a surface normal map for one or more objects in the input image, e.g. for use in an augmented reality of other application.

[0201] As some other example, a dense prediction task conditioned on an input image can be used to generate an image data item that is a version of the input image comprising more pixels (super-resolution) than the input image, or fewer pixels than the input image, or that is a de-noised version of the input image, or that is an in-painted or out-painted version of the input image.

[0202] The generated data item from such a dense prediction task may be used to provide an input to a control system of a mechanical agent, such as a robot or vehicle operating in a real -world environment, e.g. to detect real world objects to manipulate or that are obstacles, and may be used by the control system e.g. to make decisions on how to control the mechanical agent to accomplish a task performed by the robot, or for controlling the direction or speed of movement of the mechanical agent.

[0203] A dense prediction task on a (digital / digitized) audio signal may generate a data item that labels parts of the audio signal with values that are the output of the dense prediction task. The types of task performed may correspond to those described above, e.g. audio signal segmentation (semantic or instance), audio object detection (e.g. detecting particular sounds, or words, e.g. hotword detection), and so forth, generating labels rather than pixel values.

[0204] As another example a data item for a 3D image may comprise data for a spatial location and for a representation of a viewing direction, e.g. the data defining a radianceemitted in the viewing direction at the spatial location in the scene, and optionally, a volume density at the spatial location in the scene. This data may then be processed to generate a 3D image. The conditioning input can specifies the type of image to be generated, e.g. a type of object or action depicted in the still or moving image, and / or a location or viewpoint for which the image should be generated.

[0205] As another example the data item may comprise an audio data item, e.g., data defining a waveform of audio or a spectrogram (an image representing the audio), e.g., a mel-spectrogram, e.g. comprising a plurality of time or time-frequency samples. For example the audio data item may define speech in a natural language. In this example, the conditioning input can be text or features of text that the audio should represent, i.e., so that the system serves as a text-to-speech machine learning model that converts text or features of the text to audio data for an utterance of the text being spoken. Also or instead the conditioning input can identify a desired speaker for the audio, i.e., so that the system generates audio data that represents speech by the desired speaker. As another example, the conditioning input can specify a classification for the audio data into a class from a set of possible classes, so that the system generates audio data that belongs to the class. For example, the classes can represent types of musical instruments or other audio emitting devices, i.e., so that the system generates audio that is emitted by the corresponding class, types of animals, i.e., so that the system generates audio that represent noises generated by the corresponding animal, and so forth.

[0206] As another example the data item may comprise data defining a sequence of tokens, such as words or wordpieces, from a vocabulary of tokens. Thus the data item may comprise text in a natural language, e.g. that represents a processed data item or that is a translation from one to another natural language. The conditioning data may define a style or sentiment of the generated text. In still further examples the input and output data item may comprise speech, video, or time series data generally.

[0207] As another example the data item may comprise data for an image or other data item completion task in which one or more missing parts of the data item are generated or filled in by the system, conditioned on data defining the data item with the missing part(s).

[0208] As another example the data item may define actions, e.g., that can be performed by an agent interacting with an environment. For example data elements can specify positions, torques, or other control signals for the parts of a mechanical agent, or higher- level control commands. The sequence of data elements can collectively represent asequence of actions that can be performed by an agent to interact with an environment. The agent can be, e.g., a mechanical agent, e.g., a robot or a self-driving vehicle, and the environment can be a real-world environment. The actions may selected in response to one or more observations of the real-world environment, e.g. to control the mechanical agent to perform a particular task such as moving, or manipulating an object. An action to be performed may be defined by the conditioning data.

[0209] As another example the data item may define the chemical structure of a chemical, e.g. as atoms or moieties of the structure and bonds, e.g. bond angles between them, or as a sequence of tokens, e.g. in SMILES notation. The conditioning data may define a desired characteristic or property of the chemical. A method may include physically synthesizing the chemical defined by the chemical structure.

[0210] In this specification, the term "configured" is used in relation to computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are "configured" to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.

[0211] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or any combination thereof. The subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by or to control the operation of a computing device or hardware. The storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these. Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device orhardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.

[0212] The term "computing device or hardware" refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field- programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.

[0213] A computer program, also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment. A program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation ofthe computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.

[0214] In this specification, the term "engine" broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of Al and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.

[0215] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or applicationspecific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.

[0216] Computers capable of executing a computer program can be based on general- purpose microprocessors, special-purpose microprocessors, or a combination of both.They can also utilize any other type of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Typically, a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both. The essential elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the Al model, the volume of data being processed, and the desired performance and latency requirements. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high-performance computing capabilities. The system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage.

[0217] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid-state drives (SSDs), and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.

[0218] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms.The selection of input and output modalities will depend on the specific application and the desired form of user interaction.

[0219] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or J AX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.

[0220] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.

[0221] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship is established through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactionsbetween the user and the system, enabling a wide range of applications and functionalities.

[0222] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0223] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0224] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

CLAIMS1. A computer-implemented method of generating a sequence of data elements that comprises a respective continuous valued data element at each position in a sequence of positions, comprising: obtaining a current sequence of data element embeddings that comprises a respective data element embedding of each data element at a position that precedes the current position, wherein each data element embedding represents a value of the data element that is a continuous variable; processing the current sequence of data element embeddings using a neural network to generate the data element at the current position, wherein the neural network comprises a sequence of attention neural network layers and an output block, wherein each attention layer has an attention layer input for each element of the sequence of data elements and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output; wherein the output block performs operations comprising: processing the attention layer output of a last attention neural network layer before the output block to generate parameters defining a continuous distribution over possible values of a data element; and sampling from the continuous distribution to obtain the value of the data element at the current position, wherein the value of the data element at the current position is a continuous variable.

2. The method of claim 1, comprising, for each position after the first position in the sequence of positions: obtaining the current sequence of data element embeddings comprising the respective data element embedding of each data element at a respective position that precedes the current position; and processing the current sequence of data element embeddings using the neural network to generate the data element at the current position; and wherein the output block performs operations comprising sampling from the continuous distribution to obtain the value of the data element at the current position.

3. The method of claim 1, comprising for each position in the sequence of positions in parallel: obtaining the current sequence of data element embeddings comprising the respective data element embedding of each data element in the sequence of data elements; processing the current sequence of data element embeddings in parallel using the neural network to generate each data element in the sequence of data elements; and wherein the output block performs operations comprising: processing the attention layer output of a last attention neural network layer before the output block to generate, for each data element in parallel, parameters defining the continuous distribution over possible values of the data element; and sampling from the continuous distribution to obtain the value of each data element.

4. The computer-implemented method of any of claims 1-3, further comprising generating the sequence of data element embeddings by processing a current sequence of data elements using a linear layer.

5. The computer-implemented method of any of claims 1-4, wherein the continuous distribution is a Gaussian distribution.

6. The computer-implemented method of any of claims 1-5, wherein the parameters defining the continuous distribution comprise parameters of a Gaussian mixture model.

7. The computer-implemented method of claim 5 or 6, further comprising modifying a variance of the continuous distribution by a scaling factor before sampling from the continuous distribution.

8. The computer-implemented method of any one of claims 1-7, further comprising truncating the continuous distribution before sampling from the continuous distribution.

9. The computer-implemented method of any one of claims 1-8, wherein sampling from the continuous distribution to obtain the value of the data element at the current position comprises sampling a plurality of possible values for the data element at the current position; the method further comprising maintaining, whilst generating thesequence of data elements, a plurality of possible current sequences of data element embeddings, each possible current sequence of data element embeddings comprising one or more of the possible values for a data element at one or more positions in the sequence and after sampling from the continuous distribution for a final position in the sequence of positions, selecting one of the possible current sequences as the generated a sequence of data elements.

10. The computer-implemented method of any preceding claim, wherein the attention mechanism maps a query and a set of key-value pairs, each derived from the attention layer input to an output from which the attention layer output is derived, wherein the output is computed as a weighted sum of the values, weighted by a similarity function of the query to each respective key.

11. The computer-implemented method of any preceding claim, processing the current sequence of data element embeddings using the neural network to generate the data element at the current position comprises processing the current sequence of data element embeddings with the neural network conditioned on a conditioning input that defines one or more characteristics of the generated sequence of data elements.

12. The method of claim 11, wherein processing the current sequence of data element embeddings using the neural network to generate the data element at the current position comprises: i) processing the current sequence of data element embeddings with the neural network conditioned on a conditioning input that defines one or more characteristics of the generated sequence of data elements to generate parameters defining a first continuous distribution over possible values of the data element, and ii) processing the current sequence of data element embeddings with the neural network conditioned on a null conditioning input to generate parameters defining a second continuous distribution over possible values of the data element; and sampling from a combined continuous distribution defined by a combination of the first continuous distribution and of the second continuous distribution to obtain the value of the data element at the current position.

13. The computer-implemented method of any one of claims 1-12, wherein each data element represents a pixel in an image.

14. The computer-implemented method of any one of claims 1-12, wherein each data element represents an audio sample in an audio waveform.

15. A computer-implemented method of generating a data item, comprising: generating a set of latent variables by generating a sequence of data elements using the method of any one of claims 1 to 14, wherein each data element defines a value of a respective latent variable; and processing the set of latent variables using a decoder neural network to generate the data item.

16. The computer-implemented method of claim 15, wherein processing the set of latent variables using the decoder neural network to generate the data item comprises processing the set of latent variables to determine a modified set of latent variables then processing the modified set of latent variables using the decoder neural network; wherein each latent variable comprises a vector and wherein processing the set of latent variables comprises: partitioning each vector into a first part and a second part; determining a first part of a corresponding modified vector from the first part determining a second part of the corresponding modified vector by modifying the first part using a learned transformation to obtain a modified first part, and combining the modified first part with the second part.

17. The computer-implemented method of claim 16, wherein the modifying the first part using a learned transformation comprises processing the first part using a trained adapter neural network to generate the modified first part.

18. The computer-implemented method of claim 16, wherein the modifying the first part using a learned transformation comprises determining a scaling and an offset based on the first part, and applying the scaling and offset to the second part.

19. The computer-implemented method of any of claims 15-18, wherein the decoder neural network, and when dependent on claim 16 the learned transformation, has been trained by training an autoencoder comprising the decoder neural network.

20. The computer-implemented method of claim 19, wherein the autoencoder is a variational autoencoder.

21. The computer-implemented method of claim 19 or 20, wherein training the autoencoder comprises, for a plurality of training data items: processing the training data item using an encoder neural network to generate a first set of parameters defining a posterior distribution of the set of latent variables; sampling from the posterior distribution to determine sampled values of the set of latent variables; processing the sampled values of the set of latent variables using the decoder neural network to generate an output training data item; and training the encoder neural network and the decoder neural network jointly using an objective function that has a first term dependent upon a difference between the training data item and the output training data item and a second term dependent upon a difference between the posterior distribution and a prior distribution for the set of latent variables.

22. The computer-implemented method of claims 15-21, wherein the processing of the set of latent variables using the decoder neural network to generate the data item is conditioned on a conditioning input comprising conditioning data.

23. A computer-implemented method of training the neural network of any one of claims 1 to 22, comprising: obtaining an autoencoder comprising an encoder neural network and a decoder neural network; and training the neural network by, for a plurality of the training data items: processing the training data item using the encoder neural network to generate the first set of parameters defining the posterior distribution of the set of latent variables; sampling from the posterior distribution to determine training sampled values of the set of latent variables;training the neural network using the training sampled values of the set of latent variables.

24. The computer-implemented method of claim 23 wherein sampling from the posterior distribution to determine training sampled values of the set of latent variables comprises: sampling from the posterior distribution to determine an intermediate training sampled values of the set of latent variables; processing the intermediate training sampled values of the set of latent variables to determine the training sampled values of the set of latent variables, wherein each training sampled value of the set of latent variables comprises a vector and wherein processing the intermediate training sampled values of the set of latent variables: partitioning each vector into a first part and a second part; determining a first part of a corresponding modified vector from the first part; determining a second part of the corresponding modified vector by modifying the first part using a learned transformation to obtain a modified first part, and combining the modified first part with the second part.

25. The computer-implemented method of claim 23 or 24, wherein obtaining the trained autoencoder comprises, for a plurality of training data items: processing the training data item using the encoder neural network to generate a first set of parameters defining a posterior distribution of the set of latent variables; sampling from the posterior distribution to determine sampled values of the set of latent variables; processing the sampled values of the set of latent variables using the decoder neural network to generate an output training data item; and training the encoder neural network and the decoder neural network jointly using an objective function that has a first term dependent upon a difference between the training data item and the output training data item and a second term dependent upon a difference between the posterior distribution and a prior distribution for the set of latent variables.

26. The computer-implemented method of claim 23, 24, or 25, wherein each of at least some of the training data items include conditioning data for the training data item that defines one or more characteristics of the respective training data item; and wherein training the neural network using the training sampled values of the set of latent variables comprises: training the neural network whilst conditioned on a conditioning input comprising the conditioning data for the training data item used to generate training sampled values of the set of latent variables.

27. The computer-implemented method of claim 26, wherein training the neural network further comprises randomly discarding the conditioning data.

28. The computer-implemented method of any one of claims 23-27, further comprising: obtaining mask data, the mask data specifying locations to be masked in the training sampled values of the set of latent variables; and replacing, based on the mask data, values stored in the specified locations in the training sampled values with zero to obtain masked training sampled values; and wherein training the neural network using the training sampled values of the set of latent variables comprises training the neural network using the masked training sampled values of the set of latent variables.

29. The computer-implemented method of claim 28, further comprising concatenating, with a data element embedding corresponding to a latent variable, a vector with the masked training sampled values of the set of latent variables, the vector indicating one or both of i) a location of masked values, ii) a location of unmasked values.

30. The computer-implemented method of any of claims 1-14, wherein processing the current sequence of data element embeddings using a neural network comprises: processing the current sequence of data element embeddings using a normalizing flow model to generate a corresponding sequence of processed elements at respective positions to the current sequence of data element embeddings, each processed element comprising a vector with a dimensionality matching the dimensionality of the corresponding data element; andprocessing the sequence of processed elements with a decoder neural network to generate parameters defining a continuous distribution over possible values of a data element.

31. The computer-implemented method of claim 30, further comprising: receiving a prompt sequence, the prompt sequence comprising either one or more tokens representing pixel values of an image, or one or more tokens representing a sequence of text; and wherein the generated sequence of data elements represents either, respectively, a sequence of text that characterizes content of the image, or pixel values of an image characterized by the sequence of text.

32. The computer-implemented method of any one of claims 13 and 15-31, wherein the data item comprises image data, the image data specifying values for pixels of an image.

33. The computer-implemented method of any one of claims 13 and 15-31, wherein the values for the pixels define the result of a dense prediction task on an image.

34. The computer-implemented method of any one of claims 14-31, wherein the data item comprises audio data.

35. A computer-implemented method of generating image data, the image data specifying values for pixels of an image, the method comprising: generating a sequence of data elements that comprises a respective continuous valued data element at each position in a sequence of positions, comprising, for each position after a first position in the sequence of positions: obtaining a current sequence of data element embeddings that comprises a respective data element embedding of each data element at a position that precedes the current position, wherein each data element embedding represents a value of the data element that is a continuous variable; processing the current sequence of data element embeddings using a neural network to generate the data element at the current position,wherein the neural network comprises a sequence of self-attention neural network layers and an output block, wherein each self-attention layer has an attention layer input for each element of the sequence of data elements and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output; wherein the output block performs operations comprising: processing the attention layer output of a last self-attention neural network layer before the output block to generate parameters defining a continuous distribution over possible values of a data element; and sampling from the continuous distribution to obtain the value of the data element at the current position, wherein the value of the data element at the current position is a continuous variable; and processing the sequence of data elements using a trained decoder neural network to generate the image data.

36. The method of claim 35, further comprising, prior to processing the sequence of data elements using the trained decoder neural network, processing each data element of the sequence of data elements using an adapter block that modifies a second part of the data element dependent on a value of a first part of the data element.

37. A computer-implemented method of training a neural network using a sequence of continuous-valued data elements, the method comprising: processing each data element using a normalizing flow model to generate a respective training data item, each training data item comprising a vector with a dimensionality matching a dimensionality of the respective data element, thereby obtaining a sequence of training data items; and training the neural network to maximise a likelihood of a next element in the sequence of training data items.

38. The method of claim 37, wherein the continuous-valued data elements represent pixels of an image.

39. The method of claim 37 or 38, further comprising: generating noise values from a Gaussian distribution; augmenting each training data item with a respective noise value; andover the course of training the neural network, modulating a magnitude of the Gaussian distribution according to a decreasing function.

40. A computer-implemented method of generating a sequence of continuous-valued data elements, the method comprising: obtaining a neural network by training the neural network using the method of any of claims 37 to 39; and using the neural network to perform the method of any of claims 1 to 14.

41. A computer-implemented method of generating values for pixels of an image, comprising: autoregressively extending a sequence of soft tokens, each soft token comprising a vector of continuous values that represents a patch of a sequence of patches of the image, comprising: processing a current sequence of soft tokens using a neural network comprising a sequence of attention neural network layers and an output block, to generate a set of parameters of a continuous distribution, and sampling from the continuous distribution to obtain a next soft token to extend the current sequence of soft tokens; and processing each soft token of the sequence of soft tokens using a trained invertible flow model to generate a set of pixel values for each respective patch of the image, the invertible flow model having been trained to process an input vector using a learned invertible transformation to generate an output vector with the same number of dimensions.

42. The method of claim 41, wherein processing the current sequence of soft tokens using the neural network comprises: processing each soft token using one more embedding layers to generate a data item embedding; and processing the data item embeddings using the neural network.

43. The method of claim 41 or 42, comprising: receiving an initial sequence of tokens representing data items of a different modality to images; andprepending the initial sequence of tokens to an initial version of the current sequence of soft tokens.

44. The method of claim 43, further comprising adding a beginning of image token at the end of the initial sequence of tokens to denote the start of image generation.

45. The method of any of claims 41-44, further comprising: detecting a mode identification token and, in response: selecting one or a plurality of token prediction heads for generating output tokens, wherein the token prediction head include at least a first, image generation token prediction head comprising the trained invertible flow model, and a second token prediction head for processing tokens representing data items of a second, different modality to images.

46. The method of any of claims 41-45, wherein processing each soft token using the trained invertible flow model comprises reducing a dimension of the vector of continuous values representing a soft token prior to processing each soft token of the sequence of soft tokens using a trained invertible flow model.

47. The method of any of claims 41-46, wherein the pixel values comprise color or intensity values or values of a dense prediction task.

48. A system including one or more computers and one or more storage devices storing instructions that when executed by the one or more computers, cause the one or more computers to perform the method of any one of claims 1-47.

49. One or more computer storage media storing instructions that when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 47.

Citation Information

Patent Citations

  • Vector-quantized image modeling

    WO2023059699A1

Cited By

  • Large language model reasoning method, device, equipment, medium and program product

    CN121615762A

  • Anti-interference perception encryption method for unmanned aerial vehicle image

    CN122293803A