Learning representations and generating new views of data items using diffusion models

By employing a neural network system with an encoder and de-noising decoder, the method effectively generates new views of data, capturing high-level semantics and enabling efficient data manipulation, thus addressing the limitations of existing technologies.

WO2025109182A1PCT designated stage expired Publication Date: 2025-05-30DEEPMIND TECH LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/EP2024/083318
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing machine learning models struggle to effectively generate new views of data and capture high-level semantics in a way that allows for efficient and high-quality data manipulation and generation.

Method used

The use of a neural network system comprising an encoder neural network and a de-noising decoder neural network, trained using self-supervision, to generate new views of data by incrementally reducing noise in output data items at multiple time steps, allowing for the capture of key properties and semantics of data items.

Benefits of technology

This approach enables the generation of high-quality data items with reduced computational resources, allowing for efficient manipulation and control of generated data, such as generating new perspective views of 3D objects with improved accuracy and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024083318_30052025_PF_FP_ABST
    Figure EP2024083318_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Systems, methods, and program code for training an encoder neural network and de-noising decoder neural network for generating an output data item such as an image or audio Training source and target data items are obtained, representing views of an object or scene, and used to train the encoder neural network and the de-noising decoder neural network. The trained encoder neural network generates representations usable for many downstream tasks. The trained encoder neural network and de-noising decoder neural network can be used together to generate new views of objects or scenes, such as a new 3D view, given just one or a few source views.
Need to check novelty before this filing date? Find Prior Art

Description

LEARNING REPRESENTATIONS AND GENERATING NEW VIEWS OF DATAITEMS USING DIFFUSION MODELSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to GB application No. 2318012.8, filed on November 24, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND

[0002] This specification relates to processing data using machine learning models.

[0003] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.SUMMARY

[0004] This specification describes systems and methods, implemented as computer programs on one or more computers in one or more locations, for training a neural network system, and for using trained components of the system. The neural network system can be trained using self-supervision, and the trained components of the system can be used to perform a wide range of tasks, such as perception, reconstruction, and editing tasks. For example implementations of the trained system can, given one or more views of a scene, generate a new view of the scene, e.g. a view from a new perspective.

[0005] In a first aspect there is described a computer-implemented method of training a neural network system. The neural network system comprises an encoder neural network. The neural network system also comprises de-noising decoder neural network for generating an output data item by incrementally reducing a level of noise in the output data item at a plurality of time steps. The output data item can comprise, e.g. an image or audio.

[0006] The method generally involves obtaining a plurality of training data items each comprising at least one source data item and a target data item. The data item may comprise an image, audio, or other data item. The source data item and the target data item represent views of an object or scene, such as an image or audio representation of the object or scene.

[0007] At each of a plurality of training iterations the method obtains one of the training data items and processes the source data item(s) in the training data item using the encoder neural network to generate at least one latent vector representing the source data item(s). A time value is obtained for one of the time steps and a noisy version of the target data item, an embedding of the time value, and the latent vector are processed, using the de-noising decoder neural network to generate a de-noising output comprising an estimated noise data item for the time step. The estimated noise data item can be used for compensating the noise in the noisy version of the target data item.

[0008] The de-noising decoder neural network and the encoder neural network are trained by backpropagating gradients of an objective function that depends on an accuracy with which the estimated noise data item, e.g. image, estimates the noise in the noisy version of the target data item, e.g. image.

[0009] There is also described a computer-implemented method of generating an output data item, e.g. an output image, by incrementally reducing a level of noise in the output data item at a plurality of time steps. One or more characteristics of the output data item can be obtained, e.g. from a user, and used to generate the output data item so that it represents the characteristic(s). In implementations where the output data item comprises an output image, the generated output image can be a 2D or 3D image that is also defined by a target viewpoint.

[0010] There is also described a system comprising one or more computers, and one or more storage devices communicatively coupled to the one or more computers. The storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the described method.[Oil] There is further described one or more non-transitory computer storage media storing instructions that when executed by one or more computers perform the operations of the described method.

[0012] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.

[0013] Implementations of the described techniques can be used to generate representations of data items that capture their high-level semantics in a particularly useful and effective way. The described techniques can also be used to generate new examples of data items that are semantically similar to existing data items, such as new views of 2D or 3D images, e.g. new perspective views of 3D objects.

[0014] In particular, implementations of the system use a de-noising decoder neural network that enables generation of a data item based on a latent representation of the data item without needing to fully define all the characteristics of the data item. That is the described approach enables the system, in particular the latent vector, to represent just the most salient or descriptive qualities of the data item. For example, in the case of an image this may comprise the high-level semantics of the image whilst generation of localized and high frequency details are delegated to the de-noising decoder neural network.

[0015] The latent representation of the data item as one or more latent vectors provides a bottleneck between the encoder neural network and the de-noising decoder neural network that encourages the system to learn latent vector representations that capture key properties and semantics of the data items used in training, typically in an explicit and interpretable manner.

[0016] Generally, implementations of the system can be used both to generate and to modify data items such as images, audio, and other data items. A de-noising decoder neural network trained as described herein can produce very high quality output data items.

[0017] The system can learn latent vector representations that are disentangled, e.g. according to a disentanglement score, i.e. that can separate factors of variation responsible for the content of a data item, such as the appearance of an image, in particular where the data item has a real-world origin. This facilitates manipulation and control of generated data items, e.g. for generating a modified version of a data item. For example in the case of an image this can enable changing the tone, clarity or lighting conditions, or changing the characteristics of an object represented in the image such as the dimensions or material of the object or the age or gender of a person in the image.

[0018] The representations learnt by the system are also useful for downstream prediction tasks, such as data item, e.g. image, classification, and can provide improved accuracy by comparison with some other techniques.

[0019] Implementations of the described systems can be trained without needing labelled training data, using self-supervision.

[0020] The described techniques can be used to generate examples of output data items with higher quality than some other techniques, and with reduced use of computational resources. For example, in implementations good quality images can be generated with as few as 20 de- noising steps, with computational requirements that are significantly less than some other methods, e.g. NeRF-based approaches.

[0021] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] FIG. 1 show an example training system for training a neural network system comprising an encoder neural network and a de-noising decoder neural network.

[0023] FIG. 2 is a flow diagram of an example process for training a neural network system comprising an encoder neural network and a de-noising decoder neural network.

[0024] FIG. 3 illustrates different views of an object.

[0025] FIG. 4 shows an example noise schedule.

[0026] FIG. 5 illustrates types of features that can be represented by sub-vectors of a latent vector.

[0027] FIG. 6 illustrates choice of a scale factor.

[0028] FIG. 7 is a flow diagram of an example process for generating an output data item.

[0029] FIG. 8 is a flow diagram of an example process for performing a data item processing task.

[0030] FIG. 9A and 9B show examples of modified images generated by the described techniques.

[0031] FIG. 10A and 10B illustrate the generation of new 3D view of objects using the described techniques.

[0032] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0033] The described techniques generally involve an encoder neural network, and a de- noising decoder neural network. The encoder neural network is configured to encode a data item, such as an image, into one or more latent vectors that represent the data item, that are then used by the de-noising decoder neural network to guide the synthesis of a new but related data item.

[0034] The de-noising decoder neural network is configured to implement an iterative de- noising process that is conditioned on the latent vector(s), such as a diffusion model or consistency model. The encoder neural network and de-noising decoder neural network canbe trained using self-supervision, and the latent vector(s) learned like this capture visual semantics useful for reconstruction, editing and synthesis tasks as well as downstream perception tasks.

[0035] As a particular example, a neural network system comprising the encoder neural network, and the de-noising decoder neural network can implement a self-supervised diffusion model. The system uses a latent vector to provide a bottleneck between the encoder neural network, and the de-noising decoder neural network, and during training the system learns compact and useful representations.

[0036] Some implementations of the system are trained using images, but the described techniques are not limited to images. In implementations the system is trained to generate novel views of an object or scene as the self-supervised objective, which helps the system to capture visual semantics in an unsupervised manner. The representations learned like this, which typically disentangle different factors of variation in the training data, are useful for many image processing tasks, such as image reconstruction, editing, and synthesis tasks. In general all these image processing tasks involve generating an image, e.g. a reconstructed or reduced noise version of an input image, or an edited version of an input image, or a synthesized version of an input image, e.g. representing a 3D object or scene from a new pose or perspective. FIG. 10, described later, illustrates generating new 3D views of some objects.

[0037] An example system uses the image encoder to encode an input view into a latent vector, i.e. a low-dimensional representation, that is then used to guide the synthesis of a novel output view. Views can be any set of images that hold some relation among each other, visual or semantic, such as various augmentations or distortions of an original image, or different poses or perspectives of a 3D object, or simply images that share the same semantic category with one another. FIG. 3, described later, shows some examples of different views of various objects.

[0038] The de-noising decoder neural network uses the latent vector to perform image de- noising rather than pure reconstruction, which frees the encoder from having to compress all the information about the image into the representation, and instead allows it to focus on an image’s most distinctive and descriptive qualities. The encoder neural network learns to encode a source view into a latent vector that the de-noising decoder neural network uses to generate a target view that is visually or semantically related to the source view. The learned representation, i.e. the latent vector, learns to capture the most prominent commonalities between the source and target views.

[0039] There are several additional techniques that can be incorporated into the system, as described later, to improve these representations. For example, FIG. 5, inset on FIG. 1 and described in more detail later, illustrates a “layer modulation” technique that facilitates generating disentangled representations. This partitions the latent vector into sub-vectors that each modulate a respective pair of layers (the down-sampling and up-sampling layers of) a U- Net type de-noising decoder neural network, promoting specialization among the latent subvectors as indicated in FIG. 5. This can be particularly useful, for example, for image editing; for novel 3D view synthesis a cross-attention based approach as described later can be better.

[0040] FIG. 1 shows an example training system for training a neural network system. The training system of FIG. 1 can be implemented as computer programs on one or more computers in one or more locations.

[0041] The neural network system comprises an encoder neural network 120, and a denoising decoder neural network 130. In general the encoder neural network 120 and the denoising decoder neural network 130 can have any appropriate neural network architecture including, e.g., one or more feedforward neural network layers, or convolutional neural network layers, or attention neural network layers.

[0042] The encoder neural network 120 is configured to process a source data item 110, according to learnable parameters, e.g. weights, of the encoder neural network, to generate at least one latent vector, z, 122 representing the source data item. In some implementations, as described later, the encoder neural network 120 is configured to process multiple source data items 110, e.g. aggregating their respective latent vectors.

[0043] As one example, where the source data item 110 comprises an image, the encoder neural network 120 may comprise a ResNet blocks (He et al. “Deep residual learning for image recognition”, Proc. IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016). FIG. 1 shows as an example a source image that comprises a first, source view of a tiger. The encoder neural network 120 can then encode the source data item 110 as a single tZ-dimensional latent vector. As another example, the de-noising decoder neural network 130 may have a ViT (Vision Transformer) architecture (Dosovitskiy et al. arXiv:2010.11929, 2021). In some implementations the encoder neural network 120 can encode the source data item 110 as multiple latent vectors.

[0044] As another example, where the source data item 110 comprises audio data the encoder neural network 120 may comprise the audio encoder of an audio language model or of a speech recognition system such as BEST-RQ (Chine et al. arXiv:2202.01855).

[0045] In general a number of dimensions of the latent vector, z, will depend on the nature of the training data, e.g. on how many different factors of variation in the training data it is desired to represent. Merely as an illustration, d can be between, e.g. 32 and 8096.

[0046] The de-noising decoder neural network 130 is configured to process a noisy data item 132, e.g. pixel values of an image (as shown in the thumbnail), an embedding 124 of a time value, t, that identifies a de-noising time step, and the latent vector, z, 122, according to learnable parameters, e.g. weights, of the de-noising decoder neural network, to generate a de-noising output 134.

[0047] As described in more detail later, the encoder neural network and the de-noising decoder neural network are trained using a target data item, e.g. a target image. FIG. 1 shows as an example a target image that comprises a second, target view of a tiger. In implementations the de-noising decoder neural network is trained to generate the de-noising output based on an accuracy with which the de-noising decoder neural network 130 estimates the noise in a noisy version of the target data item, e.g. of the target image. The de-noising decoder neural network 130 can estimate this noise either by explicitly generating an estimate of the noise (at the de-noising output 134), or by estimating a de-noised version of the noisy data item 132 (at the de-noising output 134).

[0048] Thus in general the de-noising output 134 comprises an estimated noise data item for the time step, suitable for compensating the noise in the noisy version of the target data item i.e. suitable for use in reducing a level of noise in the noisy version of the target data item, either by explicitly estimating the noise or by estimating a reduced noise version the noisy version of the target data item. Actual use of the de-noising decoder neural network to reduce a level of noise during a data item generation, e.g. image generation, task, takes place during inference.

[0049] In general the de-noising decoder neural network 130 can be configured to implement a diffusion model or a consistency model. In general the de-noising output 134 can comprise an estimate of the noise in the noisy data item 132 or an estimate of a de-noised version of the noisy data item 132.

[0050] As one example, in inference a current noisy data item 132 at a time step can be processed to generate the de-noising output 134 in which the estimated noise data item for the time step comprises a prediction of noise that can be combined with the current, noisy data item at the time step to obtain an updated, reduced noise version of the current noisy data item for a next de-noising iteration time step. For example the estimated noise data item, inparticular a scaled version of the estimated noise data item, may be subtracted from the current noisy current data item, e.g. as described later.

[0051] As another example, in inference a current noisy data item 132 at a time step can be processed to generate the de-noising output 134 in which the estimated noise data item for the time step comprises a prediction of a de-noised version of the noisy data item at the time step. This can then be used as an updated, reduced noise version of a current noisy data item for a next de-noising iteration time step.

[0052] In implementations the de-noising decoder neural network 130 has an architecture that allows the neural network to map an input of a given dimensionality to an output of the same dimensionality, i.e. to map the noisy data item 132 to the estimated noise data item.

[0053] As an example the de-noising decoder neural network 130 may have a U-Net architecture (Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, arXiv: 1505.04597). Such an architecture can include, e.g. one or more ResNet blocks and / or one or more self-attention layers, and one or more skip connections between neural network layers of corresponding resolution. For example a U-net can be implemented as a stack of residual, convolutional, and either down-sampling or up-sampling layers (in the encoding and decoding part of the U-Net respectively), that are further linked by symmetric skip connections.

[0054] However, more generally, the de-noising decoder neural network the de-noising process can operate in either a space of the data item, e.g. in an image space, or in a latent space. Thus as some other examples the de-noising decoder neural network 130 may comprise a diffusion transformer (DiT) (Peebles et al., “Scalable Diffusion Models with Transformers” arXiv: 2212.09748, 2023), or a transformer backbone, or a U-ViT (Hoogeboom et al., arXiv: 2301.11093, 2023).

[0055] Throughout this specification, an “embedding” of an entity refers to a representation of the entity as an ordered collection of numerical values, e.g., a vector or matrix of numerical values.

[0056] In implementations the embedding of the time value, t, that identifies a de-noising time step or (later) of one or more components or coordinates of a viewpoint may be determined by encoding it as a d-dimensional vector. Any suitable encoding (embedding) may be used. As one example a sinusoidal embedding can be used in which the value of each dimension z may be sin(rnt) for even z and cos(rnt) for odd z, where a> = N~2l^dwhere / Vis a large number e.g. 10000, and where t is the time value (or coordinate of the viewpoint).

[0057] The de-noising decoder neural network 130 is configured to generate the de-noising output 134 conditioned on the embedding 124 of a time value, t, and the latent vector, z, 122. In some implementations the embedding of the time value, t, is combined with the latent vector, z, e.g. by concatenating or summing these, and the de-noising decoder neural network 130 is conditioned on the combination. In other implementations the de-noising decoder neural network 130 is conditioned on the embedding of t and latent vector z separately.

[0058] There are many ways to condition the de-noising decoder neural network 130 on the embedding of t and latent vector z. As one example this can be done by incorporating FiLM (feature-wise Linear Modulation) layers (Perez et al., arXiv: 1709.07871) into the neural network. As another example this can be done by including cross-attention blocks in the neural network with queries derived from the de-noising decoder neural network activations querying keys and values derived from the conditioning information. Some particularly useful ways of conditioning the de-noising decoder neural network 130 are described later.

[0059] In implementations where the trained system is used for image generation, the generation of the de-noising output 134 can also be conditioned on viewpoint data defining a target viewpoint for the generated image. For example, where the source data item 110 comprises a source image 110A, and the noisy data item 123 comprises a noisy image 132A, this can be done by determining a viewpoint embedding HOB, 132B that represents coordinates of a viewpoint for each pixel in the respective source image 110A and / or noisy image 132A, and processing the viewpoint embedding using, respectively, the encoder neural network 120 and / or the de-noising decoder neural network 130.

[0060] The viewpoint embeddings for the pixels of an image may be combined, e.g. pixelwise, with pixel values of the image, e.g. by concatenation or addition. Also or instead the viewpoint embeddings for the pixels may be combined with a representation of the image in the encoder neural network 120 and / or in the de-noising decoder neural network 130. For example the viewpoint embeddings for the pixels may be combined with a representation of the image at or after the output of a first neural network layer of the respective neural network (which may, but need not, preserve dimensions of the image), e.g. by concatenating a sinusoidal position embedding to linearly mapped RGB channels of the image after a first (input) layer of the encoder neural network 120 and / or after a first (input) layer of the de- noising decoder neural network 120.

[0061] The training system includes a training engine for training the encoder neural network 120 and the de-noising decoder neural network 130.

[0062] FIG. 2 is a flow diagram of an example process for training a neural network system comprising an encoder neural network and a de-noising decoder neural network, and for convenience described with reference to FIG. 1. The process of FIG. 2 may be implemented by one or more computers in one or more locations. One or both of the encoder neural network and the de-noising decoder neural network may be pre-trained; or they may be trained from scratch.

[0063] The method involves obtaining a plurality of training data items (step 200). Each training data item comprising at least one source data item 110 and a target data item. The source data item and the target data item represent views of an object or scene.

[0064] Here a “view” may comprise an observation of the object or scene, which may be a real, tangible object or scene, or an intangible object such as a data object that, in general, may represent any type of entity. Some examples of the entities that the source data item and target data item represent views of are given later.

[0065] As one example, the source data item and the target data item can represent different views of the same object, e.g. one may be a modified (“augmented”) or distorted view of the object and the other an unmodified view of the object. As another example the source data item and the target data item may comprise images that are views of the same physical object from different viewpoints. As a further example, the source data item and the target data item may represent different views of the same type or class of object, e.g. they may represent views of a cat, but not necessarily the same cat. That is, the source data item and the target data item may both represent views of the same (type of) object by may comprise different examples of the object, i.e. they may be semantically related. In some implementations, e.g. where the system is being trained for a reconstruction task, the source data item and the target data item can represent the same view, e.g. of an object or scene. By way of example, FIG. 3 illustrates, in each row, three different views of an object, illustrating image crops (top), image augmentations and distortions (middle), and camera viewpoints (bottom).

[0066] The method involves performing a plurality of training iterations.

[0067] At each iteration the process obtains, e.g. samples, one of the training data items (step 202) and processes the one or more source data items 110 in the training data item, using the encoder neural network 120, to generate the one or more latent vectors 122 representing the one or more source data items (step 204).

[0068] The process also obtains a (random) time value for one of the plurality of data generation time steps, e.g. by sampling from a distribution, e.g. from a uniform distribution (step 206). In general the time steps span a range between one end time step, e.g. at t = 1 ort = 0, i.e. with a time value of 1 or 0, and another end time step, e.g. at t = T where T is an integer. That is, in implementations the time value is an integer valued time index.

[0069] In some implementations in inference (i.e. when the de-noising decoder neural network generates an output data item by incrementally reducing a level of noise in the output data item at a succession of time steps) the time steps, i.e. time values, may count down from an initial step for which t = T to a final time step for which t = 1 or t = 0. However the direction, and which is the initial and which is the final time step, is an arbitrary choice.

[0070] In some implementations the inference process may be strided, i.e. when the denoising decoder neural network is generating the output data item an update step as described later may be performed every S time steps rather than at every time step.

[0071] The method also involves processing a noisy version of the target data item 132, an embedding of the time value 124, and the latent vector 122, using the de-noising decoder neural network 130, to generate the de-noising output 134 comprising the estimated noise data item for the time step (step 208).

[0072] In general noisy version of the target data item 132 and the estimated noise data item each have a dimension or dimensions that match the target data item (and as described later in inference, the current data item). For example if the target data item comprises an image represented by an H x V array then the noisy version of the target data item and the estimated noise data item can also each be represented by an H x V array.

[0073] In implementations the de-noising decoder neural network and the encoder neural network are trained (e.g. jointly) by backpropagating gradients of an objective function that depends on an accuracy with which the estimated noise data item estimates the noise in the noisy version of the target data item (step 210).

[0074] That is in implementations, gradients of the objective function may be backpropagated through the de-noising decoder neural network 130 and into the encoder neural network 120 to update trainable parameters, e.g. weights, of each of the neural networks.

[0075] In some implementations only part of the de-noising decoder neural network 130 and / or the encoder neural network 120 is trained, e.g. an adapter neural network part of one or both of these that is used to adapt a pre-trained part of the respective neural network for which the parameters, e.g. weights, are fixe during the training.

[0076] The training may use any appropriate gradient descent optimization algorithm, e.g. Adam or another optimization algorithm.

[0077] The following example describes one implementation of a de-noising diffusion model but the methods and systems described herein can use variants of this approach, or other diffusion or consistency model techniques.

[0078] In some implementations the noisy version of the target data item is obtained by adding (or subtracting) a noise data item that represents noise to the target data item, and a value of the objective function can then depend on a difference between the noise data item and the estimated noise data item.

[0079] For example in some implementations the noisy version of the target data item can be obtained by sampling a noise data item from a noise distribution, e.g. a Gaussian noise distribution, a mixture of Gaussian noise distributions or a Gamma noise distribution. The noisy version of the target data item can then be determined using the noise data item, e.g. using the noise data item scaled by a scale factor that depends on the time value. For example the (scaled) noise data item may be added to or subtracted from the target data item or a scaled version of the target data item, e.g. a version of the target data item scaled byIn implementations a value of the objective function can then depend on a difference between the noise data item and the estimated noise data item.

[0080] As one example if, in inference the time values count down from an initial time step for which t = T the scale factor may reduce at successive time steps. In some implementations the scale factor is defined by a value of (1 — at), where atdepends on the time value, e.g. where atis equal to or greater than zero and equal or less than 1. Then the value of atmay increase as the time value (time index) decreases, i.e. so that a1> ■■■ > at> ••• > aT. For example the scale factor may be 1 at an initial time step; the scale factor can reduce to 0 at a final time step.

[0081] The noisy version of the target data item, xt, may, e.g., be determined as xt= tx0+ (1 — at)€ where x0is the target data item and e is the noise data item, e.g. 6~JV'(0, / ). The value of the objective function can depend on a difference between the noise data item, e, and the estimated noise data item, eefrom the de-noising decoder neural network 130, e.g. on an LI or L2 norm or other difference measure. For example the value of the objective function can be determined by evaluating || e — 60||2. This can be averaged over samples from the noise distribution and over values of t which may, e.g. be sampled from a uniform distribution across values that define the time steps ({0, T}).

[0082] As illustrative technical background, in one example diffusion model implementation a diffusion process is modelled aswhere atdefines a variance schedule, and a de-noising process can involve sampling a value of xt_i according tovector as z = £(x', c') where x' is a clean (noise-free) source data item and c' is optional conditioning data such as viewpoint (embedding) data for the source data item; and the denoising decoder neural network 130 can determine the estimated noise data item as ee= D(xt, t, z, c) where c denotes optional conditioning data such as viewpoint (embedding) data for the target data item.

[0083] The conditioning c for the de-noising decoder neural network 130 can, e.g., be included with the noisy data item. In the case of an image this can be done by concatenating or otherwise combining a viewpoint embedding 132B with a noisy version of the target data item, i.e. target image, 132A.

[0084] As previously mentioned the noisy version of the target data item can be determined using, e.g. by adding, the noise data item scaled by a scale factor. In some implementations the scale factor can be defined by a value of (1 — at), where atdepends on the time value. The method can then involve determining atsuch that a gradient of the variation of atwith the time value has a vertical asymptote at a time value for an initial one of the time steps, e.g. at which t = T, and at a time value for a final one of the time steps, e.g. at which t = 1 or t = 0. More generally, in implementations a variation of the scale factor with the time value has a vertical asymptote at one or both of a time value for an initial one of the time steps and a time value for a final one of the time steps. Put differently, a gradient of a noise schedule curve representing the variation in the scale factor, or the variation in at, with time value can have a gradient that is less than unity for a range of time values around a central time value that is mid-way between the initial and final value, where the range is less than 50%, 30% or 20% of the total range. That is, a schedule of variation of atmay be arranged to emphasize training using noisy versions of the target data item with mid-range levels of noise compared with higher or lower levels of noise. Note that the manner in which at, and hence the scale factor (1 — at), varies does not depend on whether or not the target data item is scaled.

[0085] FIG. 4 shows an example noise schedule that can be used in implementations of the described training techniques. In particular FIG. 4 shows a graph of at, against time step (normalized to a range between 0 and 1).

[0086] In more detail, FIG. 4 shows a conventional noise schedule 400 (a cosine schedule) and an inverted noise schedule 410 as described above. The inverted noise schedule prioritizes medium noise levels over high or low noise levels, which can aid representation learning. For example too little noise does not present the de-noising decoder neural network 130 with a challenging enough task, reducing reliance on the encoder neural network 120; whilst too much noise can require the latent vector z to encode fine details of the generated data item, turning de-noising into mere reconstruction.

[0087] Merely as one example a set of curves similar to the inverted noise schedule 410 and parameterized by T can be obtained by inverting / [cr(e / T) =<J(S / T)] where s and e are start and end hyperparameters (to illustrate, e.g., s = — 3 or s = 0, e = 3); and < (■) is the sigmoid function. As an illustration one example curve can be obtained by setting, e.g. T = 0.5 and then by inverting, i.e. determining the inverse function of, the resulting function, in this example to obtain -0.0833*log(0.0025(-l-1.0050 / (- 1.0025+t))). The inverse function can be determined using appropriate computer software, e.g. a computer algebra system such as SymPy.

[0088] As previously mentioned, in some implementations the de-noising decoder neural network is configured for performing a de-noising process according to a diffusion model, e.g. a DDPM (Denoising Diffusion Probabilistic Model) or DDIM (Denoising Diffusion Implicit Model) model.

[0089] In some implementations the de-noising decoder neural network is configured to generate the output data item by, at each of the time steps processing a current data item, an embedding of the time value for the time step, and a latent vector, using the de-noising decoder neural network, to generate the de-noising output comprising the estimated noise data item for the time step, and updating the current data item using the estimated noise data item for the time step, to compensate for noise in the current data item, to obtain an updated version of the current data item. The output data item can then be obtained as the updated version of the current data item at a final one of the time steps. The current data item for an initial step of the process can be a random data item, e.g. sampled from a noise distribution. In these implementations the de-noising process can operate in a space of the data item, e.g. an image space.

[0090] In implementations the de-noising process uses strided time steps, i.e. rather than updating the version of the current data item at every time step, dividing the total number of time steps, T, used during training into a reduced number of steps at which the current data item is updated.

[0091] In some implementations the source data item and the target data item of a training data item may be the same data item, in which case the neural network system can be trained as an autoencoder and used for data item reconstruction. In some implementations, as previously described, the source data item and the target data item of a training data item are different.

[0092] Where the source and target data items comprise images these can comprise various augmentations or distortions of an original image, or they can show a 3D object from different poses and perspectives, or they can simply share the same semantic category with one another.

[0093] In implementations a source data item may be obtained by modifying a target data item, or vice-versa. For example obtaining the plurality of training data items can involve obtaining a plurality of target data items and, for each of the plurality of target data items generating the at least one source data item and from the target data item by modifying the target data item, or generating the target data item from the at least one source data item by modifying the at least one source data item.

[0094] As one example, a source data item may be obtained by random “augmentation” of a target data item, or vice-versa. In general one or both of a source data item and a target data item may be randomly augmented. In some implementations one or more source data items in a training data item may be augmented by adding noise to the source data item (prior to processing the source data item using the encoder neural network). This can improve performance of the trained system on later tasks.

[0095] When the source data item and the target data item comprise image data items representing images, the modifying can involve, e.g. one or more of: randomly re-sizing the image, randomly cropping the image, randomly flipping the image in space, and applying RandAugment (Cubuk et al. arXiv: 1909.13719). Random cropping may involve selecting a random patch of the image and then expanding this to the original size of the image. Flipping the image may involve applying a horizontal or vertical flip to the image.

[0096] For an image, further possibilities include one or more of: color jittering, color dropping, Gaussian blurring, solarization, rotation, masking part of the image, and an adversarial perturbation of the image. Color jittering may comprise changing one or more ofthe brightness, contrast, saturation and hue of some or all pixels of the image by a random offset. Color dropping may comprise converting the image to greyscale. Gaussian blurring may comprise applying a Gaussian blurring kernel to the image; other types of kernel may be used for other types of filtering. Solarization may comprise applying a solarizing color transform to the image; other color transforms may be used. Masking may comprise setting pixels of a random patch of the image to a uniform value, e.g. zero. Applying an adversarial perturbation may comprise applying a perturbation that increases a likelihood that the encoder neural network 120 generates an erroneous latent vector representation (e.g. determined so as to maximize an objective based on an error in performing a task using the latent vector).

[0097] For other types of data item corresponding modifications may be applied. That is, resizing, cropping, flipping, data item value jittering (local modifications), global modifications to data item values, and so forth can be applied to any type of data item. Merely as an example, in the case of an audio data item example modifications can further include modifications to the amplitude e.g. by randomly increasing or diminishing the amplitude of the audio, or modifications to the frequency characteristics of the audio, e.g. by randomly filtering the audio.

[0098] Where a training data item comprises a plurality of the source data items each of the source data items can be processed the encoder neural network to generate a respective latent vector representing the source data item. These may then be aggregated to obtain an aggregated latent vector. The aggregated latent vector, the noisy version of the target data item, and the embedding 124 of the time value may then be processed using the de-noising decoder neural network 130 to generate the de-noising output.

[0099] In some implementations aggregating the latent vectors can involve determining a mean of the latent vectors. In some implementations aggregating the latent vectors can involve processing the latent vectors using a Transformer neural network, e.g. by processing a sequence of the latent vectors to generate an output that represents an aggregation of the latent vectors.

[0100] A Transformer network may be characterized by having a succession of self-attention neural network layers. A self-attention neural network layer has an attention layer input for each element of the input and is configured to apply an attention mechanism over the attention layer input to generate an attention layer output for each element of the input.

[0101] In implementations the de-noising decoder neural network 130 comprises a succession of neural network layers. The latent vector can be partitioned into a set of sub-vectors, e.g. m or n sub-vectors (see below), and these can be used to modulate (output) activations of neurons in the neural network layers when conditioning the de-noising decoder neural network 130 on the latent vector 122.

[0102] For example, each sub-vector can be used to modulate the activations of a respective layer or, for a U-Net, of a respective pair of layers of the de-noising decoder neural network. In the case of a U-Net (or variant thereof) the pair of layers may be corresponding layers in the respective contracting and expansive paths of the U-Net, i.e. down-sampling or up- sampling layers. For example the latent vector z can be partitioned into m + 1 sub-vectors, z£, where 2m + 1 is the number of layers in the U-Net, and each sub-vector, zt, can be used to modulate a respective pair of layers (h£, h2m-t).

[0103] Since the different layers of a U-Net correspond to different resolutions of data item feature representation this can promote specialization amongst the latent sub-vectors. FIG. 5, inset on FIG. 1, illustrates the types of features that different layers of the U-Net, and hence different sub-vectors of the latent vector, z, can represent when this type of modulation is used. For example, the external, i.e., input and output layers, can represent relatively higher- frequency features corresponding to latent sub-vectors that represent, e.g. color and texture, and the internal layers can represent relatively lower-frequency features corresponding to latent sub-vectors that represent, e.g., object or scene pose or structure. Intermediate layers can represent intermediate-frequency features corresponding to latent sub-vectors that represent, e.g., object details.

[0104] Some implementations of the training involve randomly setting sub-vectors of the set of sub-vectors to zero prior to modulating the activations of the neural network layers using the set of sub-vectors, e.g. by randomly zeroing a subset of z1;m. This may be termed “layer masking”, and effectively allows two versions of the de-noising model to be implemented in the same de-noising decoder neural network, one conditional and one unconditional, and trained jointly. Optionally at inference, i.e. when de-noising, the de-noising output 134 can be modified to compensate, e.g. by adapting the output to determine e0= D(xt, t, z, c) — 2)(xt, t, 0, c), where, as before, c is optional. As an example, a rate of layer masking can be in the range 0.01 to 0.5, e.g. around 0.1.

[0105] Layer masking can mitigate the de-noising decoder neural network’s reliance on particular sub-vector dependencies, allowing the sub-vectors to decouple and specialize independently, facilitating disentangling amongst representations encoded by a sub-vector. This can facilitate image editing and style mixing as it facilitates selectively conditioning thedecoder on selected levels of granularity such as structural or positional aspects of an image, while allowing other aspects such as lighting, texture, or color palette to vary unconditionally.

[0106] Where layer masking is not used, training whilst conditioning on the latent vector z may involve so-called classifier free guidance. This can involve randomly masking out or otherwise removing the entire latent vector z from an input of the de-noising decoder neural network 130 so as to train the neural network to generate the de-noising output both with and without guidance from the conditioning data, e.g. is described, e.g., in Ho and Salimans, arXiv:2207.12598.

[0107] In some implementations modulating activations of a neural network layer using a sub-vector (when conditioning the de-noising decoder neural network 130 on the latent vector 122) involves performing an affine transformation of the activations, h, of a layer using the sub-vector.

[0108] In some implementations modulating activations of a neural network layer involves normalizing the activations and to obtain normalized activations and scaling and / or shifting the normalized activations using the sub-vector. Normalizing is optional; there are many different normalization techniques that may be used, e.g. batch normalization, layer normalization, or group normalization. In implementations the sub-vector can be processed by one or more linear layers to obtain a scale modulation sub-vector and / or a bias modulation sub-vector, that can be used to perform the respective scaling and / or shifting. Shifting can refer to adding or subtracting a value.

[0109] Merely as one example, modulating the activations at a layer can use group normalization (Wu et al., “Group normalization”, arXiv: 1803.08494, 2018) to control the scale and bias of the normalizations, by determining the modulated activations for a layer as zsGroupNorm(h) + zbwhere zsand zbare each obtained from a linear projection of z. In some implementations, z and t are combined into wsand wband used to modulate h.

[0110] In some implementations a two-stage modulation technique is used in which the activations are scaled and / or shifted using the embedding of the time value prior to scaling and / or shifting using the sub-vector. For example the embedding 124 of the time can be included by determining the modulated activations for a layer as zs(tsGroupNorm(h) + tb~) + zbwhere tsand tbare each obtained from a linear projection of the sinusoidal embedding of t.

[0111] Also or instead of globally modifying the activations of a layer using a sub-vector, the activations may be locally modified. This can involve, for one or more of the neural networklayers, processing the activations the neural network layer using a cross-attention block that updates the activations using attention over a set of n sub-vectors. There are many different types of attention mechanism that can be used. An output of the cross-attention block can provide an input to a next (subsequent) layer of the de-noising decoder neural network.

[0112] As one example the cross-attention block can be configured to apply QKV attention, by computing a similarity between a query (Q) and a set of key (K) - value (V) pairs. The set of key -value pairs can be determined from the sub-vectors and the query from the activations. The output of the cross-attention block may comprise a weighted sum of the values, weighted by a similarity function of the query to each respective key. The similarity function may comprise, e.g., a dot product, cosine similarity, or other similarity measure; the query, keys, and values may all be vectors. For example, each of a query transformation e.g. defined by a matrix WQ, a key transformation e.g. defined by a matrix WK, and a value transformation e.g. defined by a matrix Wv, may be applied to respective inputs of the cross-attention block (the sub-vector can be used for both the key and value inputs), to derive a respective query vector Q = hWQ, key vector K = znWK, and value vector V = znWvwhich are used determine an attended sequence for the output.

[0113] Updating the activations using attention over a set of n sub-vectors can involve using QKV attention in which the query vector is derived from the activations h and the key and value vectors are derived from the set of n sub -vectors.

[0114] Modifying the activations of a layer using cross-attention can perform better when generating new 3D views; using two-stage modulation can perform better for image editing, reconstruction, and representation learning.

[0115] In some implementations the encoder neural network 120 is trained with a higher learning rate than the de-noising decoder neural network 130. This can allow the encoder neural network to adapt faster to the training data than the de-noising decoder neural network, which can help the encoder neural network to guide the de-noising decoder neural network during training.

[0116] One way of implementing this to scale down the initialized weights of the encoder by a factor of k (by scaling down the standard deviation of the initialization distribution), and then allowing the training process to scale them back up by k, effectively scaling the encoder’s gradients by k. As an example k can be greater than 1 and less than 10, e.g. , k = 2.

[0117] In some implementations the training data items can be pre-processed to re-size and / or normalize them. In some implementations the output data item can be re-sized, e.g. to increase its resolution, after it has been generated.

[0118] As previously mentioned, in some implementations the training data items comprise image data items. Then the source data items, the target data items, the noise data items, and the output data item each comprise pixel values for pixels of a still or moving image. The image may be a 2D or 3D image, in color or monochrome. As defined herein an “image” includes a point cloud e.g. from a LIDAR system, a “pixel” includes a point of the point cloud, and references to a moving image or video include a time sequence of point clouds. The image may comprise an image of the real-world, e.g. captured by a camera.

[0119] In some implementations the training data items each comprise a plurality source data items each comprising a respective source image, the target data item comprises a target image, and the source images and the target image comprise respective views of the same scene. Then the source images (and optionally also one or more of the target images) can each comprise a different respective view of the scene. Then the de-noising decoder neural network 130 can be trained to enable it to generate an output data item comprising an output image that is a further different, i.e. new, view of the scene.

[0120] This can involve, for each of the pixels in each source image, determining a viewpoint embedding that represents coordinates of a viewpoint for the pixel in the respective view of the scene. The encoder neural network can then process each of the source data item(s), and the viewpoint embeddings for the pixels of the respective source image to generate the latent vector representing the source image.

[0121] In the case of a 2D image the views may comprise translations of the image, e.g. in an x- or y- direction. The coordinates of the viewpoint for a pixel may comprise x- and y- coordinates of the pixel. For example, for each H x W source image view 110A each pixel of the image may have associated 2D (x, y) coordinates that are used to determine the viewpoint embedding HOB. To determine the viewpoint embedding the coordinates can be normalized to the range [—1,1].

[0122] The system can be trained using multiple views of an object or scene, i.e. source images with different 2D coordinates, and the target image, more particularly the noisy version of the target image, can similarly be provided with associated 2D coordinates that are used to determine a viewpoint embedding for the target image, which is another view of the object or scene.

[0123] In the case of a 3D image the views may comprise perspective views of the object or scene.

[0124] As one example, the coordinates of the viewpoint for a pixel may comprise, e.g., a ray direction, d, in particular a viewing direction for the pixel that corresponds to a direction of a ray into the scene from the pixel, and a spatial location (position), o, e.g. a position of an origin of the ray. That is, the ray direction, d, may define a direction from a position of the pixel into the 3D scene (in general different pixels will have slightly different viewing directions).

[0125] As another example, the coordinates of the viewpoint for a pixel may comprise a camera pose of a real or notional camera viewing the object or scene. Such a camera pose may be represented as a single vector p that captures its position and direction in polar coordinates. As one example this vector can be determined from pt— ps, i.e. the relative camera transformation from the source image to the target image; as another example this vector can be determined by concatenating two absolute viewpoints, [pt, ps]. The vector p can be encoded using a sinusoidal embedding as previously described.

[0126] Where the coordinates of the viewpoint for a pixel comprise a ray direction there can be many ways of determining, and representing such a ray. For example in one approach the image can be considered as defined on the focal surface (focal plane) of a notional camera: the camera converts an angle of incoming light to its optical axis, to a displacement on the focal surface. In another approach the ray can, e.g., be cast onto a sphere. The directions can be represented in, e.g., polar or Cartesian coordinates.

[0127] As a particular example a ray, r, may be represented as r = (o, d). Each pixel of an image may have an associated ray position (origin) and direction, so that the (2D) image has an associated grid of ray positions and directions to represent a 3D object or scene. The image can be a source image, or a target image or, in inference, a current image (data item). The origin of a ray can correspond to a position of the pixel in a focal plane of a camera that captures an image of the object or scene.

[0128] An embedding can be generated for each of the ray positions and directions, e.g. a sinusoidal embedding, as previously described, can be used (no matter how the viewpoint is represented). To determine the viewpoint embedding the ray position and direction can be normalized to the range [—1,1].

[0129] For example, the position may be defined coordinates (x, y) (e.g. where z is fixed) or (x, y, z), e.g. normalized in a range [+1, —1], and the direction e.g. defined as a two-dimensional vector (0, ) in a spherical coordinate system (where 0° < 0 < 180° and 0° < < 360°) or defined as a three-dimensional unit vector d. Thus, for example, for an H x W image each pixel of the image may have associated 3D coordinates (with 4, 5, or 6 dimensions) that are used to determine the corresponding viewpoint embedding. In some implementations the ray position and direction can be concatenated, e.g. as [o, d]. In some implementations the ray position and direction can be combined as a parametric sum, e.g. o + sdd, where sdis a scaling factor hyperparameter. The scaling factor, sd, can be chosen in various ways, e.g. empirically, or so as to normalize d to a unit length, or with a value determined by casting d onto the image plane or a sphere centered at the object / scene.

[0130] Where source images are images of the real world captured using a camera the ray position and direction for each pixel can be determined based on so-called intrinsic parameters of the camera, using known techniques.

[0131] In general the intrinsic parameters define a mapping from a coordinate frame of the image sensor to pixel coordinates in an image captured by the image sensor. For example the intrinsic parameters can define an intrinsic transformation matrix that transforms points from the image sensor coordinate frame to a pixel coordinate frame. Examples of intrinsic parameters include focal length, aperture, field-of-view, and resolution. Broadly the camera converts an angle (to the camera optical axis) of incoming light to a displacement on the focal surface. For example, given a camera pose p, a closed form calculation can be used to derive a 2D grid of rays r comprising ray origins and directions. Standard software packages, such as COLMAP, can be used for this.

[0132] Where a sinusoidal position embedding is used, for either 2D or 3D coordinates, optionally the arguments of sin(rnt) and cos(rnt) can be scaled by a scale factor, e.g. 2ns where s is a hyperparameter that represents the scale factor. The scale factor is chosen so as to increase the distinctiveness between different positional embeddings, i.e. so that different positions are associated with different embeddings. FIG. 6 illustrates choice of such a scale factor, representing position in the range [—1,1] on the x-axis and frequency component on the y-axis, for both sin and cos components of the embedding, illustrating a too-high scale factor (left), a satisfactory scale factor (middle), and a too-low scale factor (right). Where a scale factor is used its value can be determined empirically.

[0133] In general the viewpoint coordinates can be provided for the source image(s), for the target image, or for both. The viewpoint may be an absolute viewpoint, or a relative viewpoint of a source image relative to the target image or vice-versa. The viewpointembedding may be any embedding of the coordinates of the viewpoint, e.g. a sinusoidal embedding as previously described.

[0134] As previously described, some implementations of the training technique involve determining the viewpoint embedding for each of the pixels in the target image, processing the viewpoint embeddings for the pixels of the source image, the noisy version of the target data item (including the embeddings), the embedding of the time value, and the latent vector using the de-noising decoder neural network to generate the de-noising output.

[0135] The training can involve randomly masking, e.g. setting to zero, the viewpoint embeddings prior to processing using the encoder neural network 120, and when these are processed by the de-noising decoder neural network, prior to processing using the de-noising decoder neural network. This can help the system to use only partial information in inference.

[0136] For example, the trained de-noising decoder neural network can be used to generate an image, e.g. of an unspecified object, at a requested pose, or to generate a defined image, e.g. of a requested object, with an arbitrary pose (view). That is, at inference information such as a latent vector can be used to specify an object in the image, and / or pose information can be used to specify the pose of an object.

[0137] A training process as described above can use unlabeled training data. Each training data item need only comprise one or more source data items and a target data item, e.g. a source image and a target image, as described above. There is a wide range of publicly available datasets that can be used for training the neural network system. In general the training data items can correspond to the types of data item that, after training, will be processed and / or generated.

[0138] As some illustrative examples, datasets that can be used include: ImageNet (objects); CelebA-HQ (Liu et al. 2018; people); ShapeNet (Chang et al. arXiv: 1512.03012, 2015; 3D objects); Google Scanned Objects (GSO; Downs et al. arXiv:2204.11918, 2022; 3D objects); Co3D (Reizenstein et al. “Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction”, arXiv:2109.00512, https: / / ai.meta.com / datasets / co3d- dataset / ; 3D multi-view images of common real-world objects); Caltect-UCSD Birds (CUB- 200-2011; Welinder P. et al. "Caltech-UCSD Birds 200", California Institute of Technology, CNS-TR-2010-001, 2010, 2011; useful for evaluating disentanglement).

[0139] In addition, for generating new 2D or 3D views of an object or scene source and target images can be captured simply by capturing a video as a camera is moved around an interior or outside environment, say an office or building, logging camera location, andknowing the camera intrinsic parameters, in particular optical center and focal length. Once trained a new 2D or 2D image of the environment can be generated.

[0140] Once trained just the encoder neural network, or just the de-noising decoder neural network, or both, may be used to perform a task. Some examples of the visual and other tasks that may be performed are described later.

[0141] FIG. 7 is a flow diagram of an example process for generating an output data item. The process generally operates by incrementally reducing a level of noise in the output data item at a plurality of time steps. The processes uses a de-noising decoder neural network; for convenience the process is described with reference to the de-noising decoder neural network 130 of FIG. 1. The process of FIG. 7 may be implemented by one or more computers in one or more locations.

[0142] In implementations the process comprises obtaining an initial version of a current data item (step 700), e.g. by sampling from a noise distribution. The initial version of the current data item may comprise purely noise.

[0143] At each of a succession of time steps (which may be strided time steps) the current data item and an embedding of a time value for the time step are processed using a de-noising decoder neural network, e.g. the de-noising decoder neural network 130, after training as described above, to generate the de-noising output 134 (step 702). The de-noising output 134 comprises an estimated noise data item for the time step.

[0144] Then the current data item can be updated using the estimated noise data item for the time step (step 704). As one example the current data item can be updated using the estimated noise data item to compensate for noise in the current data item, e.g. by using the estimated noise data item as the current data item. As another example the current data item can be updated by, e.g., subtracting the estimated noise data item from the current data item (where the estimated noise data item comprises an estimate of noise in the current data item). One or both of the current data item and the estimated noise data item may be scaled. In some cases, e.g. with a non-deterministic mapping such as DDPM, noise may be added to the updated current data item (in effect, sampling from a distribution with a mean dependent on the estimated noise data item). In this way an updated version of the current data item is obtained.

[0145] In general the update step may be performed according to known diffusion model or consistency model techniques, e.g. a DDPM or DDIM update may be used. Merely as an example, a DDPM update can involve determining value ofas described above.

[0146] The output data item can comprise (be) the updated version of the current data item at a final time step (step 706). In general the final time step is after a predetermined number of de-noising time steps, as previously described.

[0147] In some implementations the output data item is generated unconditionally, e.g. by omitting or masking, e.g. setting to zero, a latent vector input of the de-noising decoder neural network. In some implementations the output data item is generated conditionally, e.g. based on a latent vector, or partially masked latent vector, provided to the de-noising decoder neural network.

[0148] Some implementations of the method involve obtaining a latent vector, e.g. the latent vector z described above, representing one or more characteristics of the output data item. This is then processed by the de-noising decoder neural network, to generate an output data item having the specified characteristics.

[0149] As one example obtaining the latent vector can involve obtaining a data item, e.g. an image, and processing the data item using the encoder neural network 120 to generate the latent vector z representing the data item. One or more elements of the latent vector representing the data item may then be modified to obtain a latent vector for processing by the de-noising decoder neural network 130.

[0150] When trained as described above the encoder neural network can generate disentangled latent variable representations of data items, which facilitates modifying aspects of the output data item by modifying the latent vector processed by the de-noising decoder neural network.

[0151] In some implementations, in particular where the training has involved layer masking, generating the output data item can involve processing the current data item and the embedding of a time value for the time step, and a masked version of the latent vector, using the de-noising decoder neural network to generate a second de-noising output comprising a second estimated noise data item for the time step. The current data item can then be updated using a combination of the estimated noise data item for the time step and the second estimated noise data item for the time step. For example the current data item can then be updated using a difference between the estimated noise data item and the second estimated noise data item, e.g. to determine e0= D(xt, t, z, c) — D(xt, t, 0, c) as described above.

[0152] As previously described, in some applications the current data item comprises a current image, the output data item comprises an output image, and generating the output dataitem comprises determining values for pixels of the current image and for pixels of the output image.

[0153] In some applications, the latent vector can represent a compressed version of an input image. For example, the latent vector representing the compressed version of the input image may have been obtained by processing the input image using the encoder neural network, to generate the latent vector. Such a compressed latent vector may, e.g. be stored or transmitted, e.g. over a network. Generating the output data item comprising the output image can reconstruct a version of the input image, e.g. representing corresponding semantic content.

[0154] When the output data item comprises an image, generating the image may involve obtaining viewpoint data defining a target viewpoint for the output image, determining a pixel viewpoint for each of the pixels in the current image from the target viewpoint, and, for each of the pixels in the current image determining a viewpoint embedding for the pixel viewpoint. The de-noising decoder neural network 130 can then process the values for the pixels of the current image, the viewpoint embeddings for the pixels of the current image, the embedding of a time value for the time step and (optionally) the latent vector, to generate the output image. The pixel viewpoints may be obtained as described above; the viewpoint embeddings may be determined and processed as described above.

[0155] As one example, for a new 2D image the target viewpoint (view) for the output image may define a 2D position for the image and the pixel viewpoint for each of the pixels may comprise an (x, y) coordinate for each of the pixels.

[0156] As another example, for a new 3D image the target viewpoint (view) for the output image may define a pose or orientation of a notional camera capturing the output image. The pixel viewpoint for each of the pixels may comprise a ray as r = (o, d) for each pixel, or a camera pose vector vector p for each pixel, as previously described.

[0157] In general generating an new view of an output data item, such as an image of an object or scene, involves obtaining one or more source data item, e.g. image(s), and corresponding viewpoint embeddings, e.g. for the pixels of the source image(s), that represent the respective views of the source data item(s). The or each source data item, e.g. image, is processed by the encoder neural network 120 to generate the latent vector z representing the data item. Where there are multiple source data items, the latent vectors may be aggregated as previously described. The (aggregated) latent vector is then used to condition generation of the new data item, e.g. new image, having the new viewpoint (view) by the trained denoising decoder neural network 130.

[0158] FIG. 8 is a flow diagram of an example process for performing a data item processing task; for convenience the process is described with reference to the encoder neural network 120 of FIG. 1. The process of FIG. 8 may be implemented by one or more computers in one or more locations.

[0159] The process involves receiving a latent vector from the encoder neural network 120, in particular after training as described above (step 800), and using the latent vector to perform the data item processing task (step 802).

[0160] As one example the data item processing task may comprise an image processing task. Then the data item comprises an image, the latent vector has been generated by the encoder neural network by processing pixels of the image, and performing the image processing task can comprise using the latent vector, e.g., to predict one or more characteristics of the image, e.g. in a classification task or other prediction task to predict a class or category or property (e.g. color / palette, shape, texture, pose, lighting, semantic attribute etc.) of the image or of one or more objects in the image.

[0161] In another example the data item processing task may comprise a control task. For example the data item may comprise an image, the latent vector has been generated by the encoder neural network by processing pixels of the image, and performing the data item processing task may comprise using the latent vector to control an action of a mechanical agent acting in a real-world environment to perform a mechanical task.

[0162] A description of a few example uses of the techniques and neural networks described herein for data item generation and editing follows.

[0163] The above described method of generating an output data item may be used to generate any type of data item, unconditionally or conditioned on a latent vector (z) that defines characteristics of the data item. Because the latent vector representation are generally disentangled, values for different elements of the latent vector can represent semantically meaningful factors of variation of the data item, and may be chosen to define selected characteristics of the data item.

[0164] As described above, the data item may comprise an image and the latent vector can define, e.g., an object and / or characteristics of an object, to be represented by the image. Where the image is a moving image, this may be generated to represent a characteristic motion of an object in the image, e.g. motion of a vehicle on a road or human motion such as walking or running.

[0165] As another example, the data item may comprise an audio data item representing an audio signal, e.g. as a values for a digital waveform of the audio signal or as a spectrogram,e.g., a mel-spectrogram. The latent vector can define a content of the audio signal, e.g. it may define words in a natural language to be represented as speech by the audio signal, and / or other characteristics such as sentiment, speaker characteristics such as age or gender, and so forth. As another example the latent vector can define a notes or other sounds to be generated by a musical instrument and / or the type of musical instrument.

[0166] As another example, the data item may represent one or more chemical molecules such as one or more proteins or ligands, e.g. as a point cloud. The latent vector can define one or more characteristics of the chemical molecule(s), e.g. in terms of its / their physical or chemical structure or properties. The data item may be used to determine a 3D structure of the chemical molecule(s), e.g. to identify one or more binding sites of or for a ligand such as a drug. This may be used as part of a screening process to identify one chemical molecule that binds to another. Such a screening process may involve evaluating an interaction of one or more candidate ligands with the structure of a target, e.g. a target protein, and then selecting one or more of the candidate ligands dependent on a result of the evaluation. For example the target may comprise a receptor or enzyme, and the ligand may be an agonist or antagonist of the receptor or enzyme. The ligand may be a drug or a ligand of an industrial enzyme. Such a process may also involve synthesizing a molecule identified by the screening process, e.g. the ligand, and optionally also testing activity, e.g. biological activity, of the molecule, e.g. ligand, in vitro and / or in vivo.

[0167] As another example, the data item may represent the output of a scientific or medical instrument, e.g. the output of an electrocardiograph or of a body scanner such as an MRI machine. The latent vector can define one or more characteristics of the data item such as whether it represents a signal from a normal body or from a diseased body. The data item may then be compared with a corresponding data item obtained from a patient, and ta comparison made to identify the likely presence or absence of a disease.

[0168] In any of the above applications the latent vector may be, but need not be, obtained by processing a data item of a similar type to the output data item using the trained encoder neural network 120, to obtain a latent vector that can then be modified according to the desired characteristics. For example the values some elements of the latent vector, e.g. those corresponding to desired properties or characteristics of the data item to be generated, may be retained, and the values of other elements of the latent vector may be modified.

[0169] In more detail, implementations of the described techniques train the encoder neural network 120 to generate a latent vector (z) that comprises human-interpretable factors of variation, i.e. values of elements of the vector correspond to particular characteristics of thegenerated data item. For example in the case of an image, the factors of variation can correspond to physical aspects of the image, such as viewing angle or lighting, or semantic aspects of the image, such as maturity (e.g. kitten vs cat).

[0170] A data item generated as described previously may be a reconstructed or edited version of a data item encoded by the (trained) encoder neural network 120, with values of one or more elements of the latent vector representation of the data item modified.

[0171] In the case of an image data item this may be used to modify a semantic content of the image, and / or a view, e.g. a perspective view, represented by the image, and / or attributes of one or more objects represented in the image such as shape, color, pose, and so forth.

[0172] In the case of an image data item this may similarly be used to modify a semantic content of the audio, and / or characteristics of the audio such as pitch, or tone color, and / or attributes of one or more objects represented in the audio, and so forth.

[0173] Where the data item represents one or more chemical molecules this may be used to modify physical, e.g. structural, or chemical characteristics of the molecule(s), e.g. to increase or decrease the interaction of the molecule(s) with another molecule or atom.

[0174] FIG. 9 shows examples of modified images generated by the de-noising decoder neural network 130 when trained and used as described above.

[0175] In particular FIG. 9A shows examples of images generated by linearly traversing the latent space from one latent vector zxto another z2, according to zx+ (z2, — zx). t for t E [0,1], illustrating a smooth transition between semantic attributes.

[0176] FIG. 9B shows examples of images generated by varying elements of the latent vector z after training on a set of N images, illustrating that these elements correspond to meaningful, disentangled factors of variation, in the illustration maturity, expression, and color. Optionally the most meaningful elements can be found by principal component analysis decomposition over the latent vectors z^=1to identify the latent directions Sj of greatest variation; in the figure examples of these are traversed according to z + Sjt.

[0177] FIG. 10 illustrates the generation of a new 3D view of an object from a requested camera perspective given one or more viewpoint-conditional source views, using the above descried techniques. In general just 1-10 source views, e.g. 1-3 source views, are needed. FIG. 10A shows examples of new 3D perspective views of an object generated from two different source image views. FIG. 10B shows examples of new 3D perspective views of various objects each generated from a single source image view.

[0178] A description of a few example uses of the representations generated using the techniques described herein follows.

[0179] The trained encoder neural network 120 may be used to process a data item to generate a representation of the data item as a latent vector that may then be used in a subsequent processing task. There are many known techniques and systems for processing a representation of a data item to perform a particular task; the trained encoder neural network 120 may be used to provide a front end, to pre-process data for any of these techni ques / sy stem s .

[0180] Where the data item comprises an image, the latent vector representation of the image may be subsequently processed to classify the image, or one or more objects in the image, or to identify the presence of one or more objects in the image, e.g. by determining a score for each category of a set of possible classification categories for the image. Where the image is a moving image this may correspondingly be used to identify the presence of one or more actions in the image or to classify an action in the image, or to predict an action or event in the image. Other tasks may be performed in a similar manner, e.g. a 3D pose estimation task to estimate the pose of one or more objects represented in an image, a shape estimation task, a task that involves identifying aspects of an image using color, a counting task that involves counting objects or objects of a particular type, or a task that involves understanding spatial relationships between objects or object attributes. In general the output of a system for performing such a task may comprise any continuous or discrete representation of the desired output data.

[0181] As another example, where the data item comprises an image that is an observation of a real-world environment, e.g. captured by a camera, the latent vector representation of the image may be used to provide an input to a control system of a mechanical agent, such as a robot or vehicle operating in the real-world environment. The latent vector can provide a compact representation in terms of factors of variation that are relevant to operation of the mechanical agent, facilitating agent control. For example the latent vector can represent in a compact form relevant objects in and aspects of the real -world environment. Thus the latent vector can be used by the control system to make decisions and / or take actions in the environment, e.g. to accomplish a task performed by the agent, or for controlling the direction and / or speed of movement of the agent.

[0182] Where the data item comprises an audio data item, e.g. a digitized audio waveform of spectrogram representation of an audio signal, the latent vector representation of the audio may be subsequently processed to classify the audio, or one or more sound objects in theaudio, or to identify the presence of one or more sound objects in the audio. As some examples the system may perform an identification or classification task such as a speech or sound recognition task, a sound or speaker classification task, an audio tagging task (in which case the output may be a category score or tag for a data item or for a segment of the data item), or a similarity determination task e.g. an audio copy detection or search task, in which case the output may be a similarity score.

[0183] As a further example, the system for which the trained encoder neural network 120 provides a front end can be a multimodal machine learning model, e.g. a visual language model. The trained encoder neural network can process an image data item or an audio data item to provide a latent vector representation for further processing by such a model.

[0184] The model output from such a model may comprise any form of output appropriate to the machine learning task performed by the multimodal machine learning model. For example the model output may comprise text in a natural or computer language that defines a result of the task, e.g. for tasks such as image captioning, visual question answering, or object detection or instance segmentation. Also or instead the model output may comprise data defining an image, video or audio object, e.g. in a generative task; or the model output may comprise non-textual action selection data for selecting an action to be performed by an agent controlled by the model.

[0185] Some example multimodal machine learning models in which the trained encoder neural network 120 may be used include: Flamingo (Alayrac et al. arXiv:2204.14198); ALIGN (Jia et al., arXiv:2102.05918); PaLI (Chen et al. arXiv:2209.06794); and PaLLX (Chen et al. arXiv:2305.18565). Some examples of multimodal machine learning models controlling an agent, in which the trained encoder neural network may be used, are described in: PaLM-E (Driess et al. arXiv:2303.03378); RT-1 (Brohan et al. arXiv:2212.06817); and RT-2 (Brohan et al. arXiv:2307.15818).

[0186] As an example of the quality of the representations learned by the trained encoder neural network 120, a linear probe classification task was performed after training on ImageNetlK (1000 object classes), fitting a linear classifier over the latent vectors z to predict object class. The described techniques reach 72% accuracy (topi prediction), a high score for representations learned using unsupervised learning.

[0187] In this specification, the term "configured" is used in relation to computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling itto carry out those operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are "configured" to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.

[0188] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or any combination thereof. The subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non- transitory storage medium for execution by or to control the operation of a computing device or hardware. The storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these. Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.

[0189] The term "computing device or hardware" refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, isemployed. Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.

[0190] A computer program, also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment. A program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.

[0191] In this specification, the term "engine" broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of Al and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.

[0192] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to performfunctions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.

[0193] Computers capable of executing a computer program can be based on general-purpose microprocessors, special -purpose microprocessors, or a combination of both. They can also utilize any other type of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Typically, a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both. The essential elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the Al model, the volume of data being processed, and the desired performance and latency requirements. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high-performance computing capabilities. The system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage.

[0194] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid-state drives (SSDs), and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used willdepend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.

[0195] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms. The selection of input and output modalities will depend on the specific application and the desired form of user interaction.

[0196] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.

[0197] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.

[0198] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship is established through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities.

[0199] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0200] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0201] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

CLAIMS1. A computer-implemented method of training a neural network system, the neural network system comprising an encoder neural network, and a de-noising decoder neural network for generating an output data item by incrementally reducing a level of noise in the output data item at a plurality of time steps, the method comprising: obtaining a plurality of training data items each comprising at least one source data item and a target data item, wherein the source data item and the target data item represent views of an object; and, for each of a plurality of training iterations: obtaining one of the training data items; processing the at least one source data item in the training data item using the encoder neural network to generate at least one latent vector representing the at least one source data item; obtaining a time value for one of the time steps; processing a noisy version of the target data item, an embedding of the time value, and the latent vector using the de-noising decoder neural network to generate a de-noising output comprising an estimated noise data item for the time step, for compensating the noise in the noisy version of the target data item; and training the de-noising decoder neural network and the encoder neural network by backpropagating gradients of an objective function that depends on an accuracy with which the estimated noise data item estimates the noise in the noisy version of the target data item.

2. The method of claim 1, further comprising: sampling the noise data item from a noise distribution; and determining the noisy version of the target data item using the noise data item; scaled by a scale factor that depends on the time value; and wherein the objective function that depends on a difference between the noise data item and the estimated noise data item.

3. The method of claim 2, wherein a variation of the scale factor with the time value has a vertical asymptote at one or both of a time value for an initial one of the time steps and a time value for a final one of the time steps.

4. The method of claim 1, 2, or 3, wherein the de-noising decoder neural network is configured to generate the output data item by, at each of the time steps: processing a current data item, an embedding of the time value for the time step, and a latent vector, using the de-noising decoder neural network, to generate the de-noising output comprising the estimated noise data item for the time step, and updating the current data item using the estimated noise data item for the time step, to compensate for noise in the current data item, to obtain an updated version of the current data item; wherein the output data item comprises the updated version of the current data item at a final one of the time steps.

5. The method of any of claims 1-4, wherein obtaining the plurality of training data items comprises: obtaining a plurality of target data items and, for each of the plurality of target data items: generating the at least one source data item and from the target data item by modifying the target data item, or generating the target data item from the at least one source data item by modifying the at least one source data item.

6. The method of any of claims 1-5, further comprising, prior to processing the at least one source data item in the training data item using the encoder neural network, adding noise to the at least one source data item.

7. The method of any of claims 1-6, wherein one or more of the training data items comprises a plurality of the source data items, the method further comprising: processing the plurality of source data items using the encoder neural network to generate, for each of the source data items, a respective latent vector representing the source data item; aggregating the latent vectors to obtain an aggregated latent vector; and processing the noisy version of the target data item, the embedding of the time value, and the aggregated latent vector using the de-noising decoder neural network to generate the de-noising output.

8. The method of any of claims 1-7, wherein the de-noising decoder neural network comprises a succession of neural network layers, and wherein processing the noisy version of the target data item, the embedding of the time value, and latent vector using the de-noising decoder neural network to generate the de-noising output comprises: partitioning the latent vector into a set of sub-vectors; and modulating activations of neurons in the neural network layers using the set of subvectors.

9. The method of claim 8, wherein modulating activations of the neural network layers using the set of sub-vectors comprises, for one or more of the neural network layers, modulating the activations the neural network layer using one of the sub-vectors by normalizing the activations and to obtain normalized activations and scaling and / or shifting the normalized activations using the sub-vector.

10. The method of claim 9, comprising scaling and / or shifting the normalized activations using the embedding of the time value prior to scaling and / or shifting the normalized activations using the sub-vector.

11. The method of any of claims 8-10, further comprising randomly setting sub-vectors of the set of sub-vectors to zero prior to modulating the activations of the neural network layers using the set of sub-vectors.

12. The method of claim 8, wherein modulating activations of the neural network layers using the set of sub-vectors comprises, for one or more of the neural network layers, processing the activations the neural network layer using a cross-attention block that updates the activations using attention over the set of sub-vectors.

13. The method of any of claims 1-12, wherein training the de-noising decoder neural network and the encoder neural network by backpropagating gradients of the objective function comprises training the encoder neural network with a higher learning rate than the de-noising decoder neural network.

14. The method of any of claims 1-13, wherein the training data items comprise image data items, and wherein the at least one source data item, the target data item, and the output data item each comprise pixel values for pixels of an image.

15. The method of claim 14, wherein one or more of the training data items comprises a plurality of the source data items each comprising a respective source image, wherein the target data item comprises a target image, wherein the source images and the target image comprise respective views of the same scene, wherein at least the source images, and wherein the de-noising decoder neural network is trained to enable the de-noising decoder neural network to generate an output data item comprising an output image that is a further different view of the same scene; the method further comprising: for each of the pixels in each source image determining a viewpoint embedding that represents coordinates of a viewpoint for the pixel in the respective view of the scene; and wherein processing the at least one source data item in the training data item using the encoder neural network comprises processing each source image in the training data item, and the viewpoint embeddings for the pixels of the source image, using the encoder neural network, to generate the latent vector representing the source image.

16. The method of claim 15, further comprising determining the viewpoint embedding for each of the pixels in the target image; and wherein processing the noisy version of the target data item, the embedding of the time value, and the latent vector using the de-noising decoder neural network to generate the de-noising output comprises processing the viewpoint embeddings for the pixels of the source image, the noisy version of the target data item, the embedding of the time value, and the latent vector using the de-noising decoder neural network to generate the de-noising output.

17. The method of claim 15 or 16, wherein the source image and the target images comprise 3D images of the scene; and wherein the coordinates of the viewpoint for the pixel in the respective view of the scene comprise either i) a spatial location and a viewing direction that corresponds to a direction of a ray into the scene from the pixel or ii) a vector that represents a camera pose of a camera viewing the scene.

18. The method of any of claims 15-17, further comprising randomly masking the viewpoint embeddings prior to processing using the encoder neural network or, when dependent on claim 16, prior to processing using the de-noising decoder neural network.

19. A computer-implemented method of training a neural network system, the neural network system comprising an encoder neural network for encoding an image, and a denoising decoder neural network for generating an output image by incrementally reducing a level of noise in the output image at a plurality of time steps, the method comprising: obtaining a plurality of training images each comprising at least one source image and a target image, wherein the source image and the target image represent views of an object; and, for each of a plurality of training iterations: obtaining one of the training images; processing the at least one source image in the training image using the encoder neural network to generate at least one latent vector representing the at least one source image; obtaining a time value for one of the time steps; processing a noisy version of the target image, an embedding of the time value, and the latent vector using the de-noising decoder neural network to generate a de-noising output comprising an estimated noise image for the time step, for compensating the noise in the noisy version of the target image; and training the de-noising decoder neural network and the encoder neural network by backpropagating gradients of an objective function that depends on an accuracy with which the estimated noise image estimates the noise in the noisy version of the target image.

20. A computer-implemented method of generating an output data item by incrementally reducing a level of noise in the output data item at a plurality of time steps, the method comprising: obtaining an initial version of a current data item and, at each of a succession of time steps: processing the current data item and an embedding of a time value for the time step using a de-noising decoder neural network to generate a de-noising output comprising anestimated noise data item for the time step, the de-noising decoder neural network having been trained by performing the respective operations of the method of any one of claims 1- 18, and updating the current data item using the estimated noise data item for the time step, to compensate for noise in the current data item, to obtain an updated version of the current data item; wherein the output data item comprises the updated version of the current data item at a final time step.

21. The method of claim 20, wherein obtaining the initial version of a current data item comprises sampling the initial version of the current data item from a noise distribution.

22. The method of claim 20 or 21, further comprising: obtaining a latent vector representing one or more characteristics of the output data item; and processing the current data item and the embedding of a time value for the time step, and the latent vector, using the de-noising decoder neural network to generate the de-noising output.

23. The method of any of claims 20-22, further comprising: processing the current data item and the embedding of a time value for the time step, and a masked version of the latent vector, using the de-noising decoder neural network to generate a second de-noising output comprising a second estimated noise data item for the time step; and updating the current data item using a combination of the estimated noise data item for the time step and of the second estimated noise data item for the time step, to compensate for noise in the current data item.

24. The method of any of claims 20-23, wherein the current data item comprises a current image, wherein the output data item comprises an output image, and wherein generating the output data item comprises determining values for pixels of the current image and for pixels of the output image.

25. The method of claim 24 when dependent on claim 22, wherein the latent vector representing one or more characteristics of the output data item comprises a latent vector representing a compressed version of an input image, and wherein generating the output data item comprising the output image reconstructs a version of the input image.

26. The method of claim 24 or 25, further comprising: obtaining viewpoint data defining a target viewpoint for the output image; determining a pixel viewpoint for each of the pixels in the current image from the target viewpoint; for each of the pixels in the current image determining a viewpoint embedding for the pixel viewpoint; and processing the values for the pixels of the current image, the viewpoint embeddings for the pixels of the current image, the embedding of a time value for the time step and, when dependent on claim 22, the latent vector, using the de-noising decoder neural network, to generate the de-noising output.

27. A computer-implemented method of performing a data item processing task, the method comprising: receiving a latent vector from an encoder neural network, the encoder neural network having been trained by performing the respective operations of the method of any one of claims 1-19; and using the latent vector to perform the data item processing task.

28. The method of claim 27, wherein the data item comprises an image, wherein the latent vector has been generated by the encoder neural network by processing pixels of the image, and wherein performing the data item processing task comprises using the latent vector to predict one or more characteristics of the image.

29. The method of claim 27, wherein the data item comprises an image, wherein the latent vector has been generated by the encoder neural network by processing pixels of the image, wherein performing the data item processing task comprises using the latent vector to control an action of a mechanical agent acting in a real-world environment to perform a mechanical task.

30. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective method of any one of claims 1-29.

31. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1-29.

Citation Information

Cited By

  • Fault diagnosis method based on V-DIT data enhancement and CNN

    CN120804833A