Methods for 3d-aware image processing
By using neural assets with appearance and pose tokens to train an image generation model, the method disentangles object appearance and pose, enabling accurate and realistic rendering of complex multi-object scenes.
Patent Information
- Application Number
- PCT/US2025/030387
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-05-21
- Publication Date
- 2025-11-27
AI Technical Summary
Existing image processing methods struggle to accurately edit or generate complex multi-object real-world scenes, as they fail to disentangle object appearance and pose, leading to inaccurate rendering and processing.
A method involving neural assets with appearance and pose tokens is used to train an image generation machine learning model, allowing for the disentanglement of object appearance and pose, enabling accurate tracking and rendering of multiple objects and backgrounds across time and changes in pose.
This approach enables the generation of highly realistic images with precise control over object and background appearance and pose, facilitating complex multi-object scene editing and generation.
Smart Images

Figure US2025030387_27112025_PF_FP_ABST
Abstract
Description
[0001]METHODS FOR 3D-AWARE IMAGE PROCESSING BACKGROUND This specification relates to image processing. The techniques have particular but not exclusive application to editing an image that takes into account the 3D appearance and pose of an object and / or background in the image. The image processing method is performed, at least in part, using a machine learning model, for example an image generation machine learning model. SUMMARY This specification describes methods, implemented as computer programs on one or more computers in one or more locations, and corresponding systems that enable image processing in the form of the generation of an edited image depicting one or more objects and / or a background. The object(s) and / or background in the edited image generated using the methods described herein each have a particular appearance (e.g. visual appearance including color, pattern, intensity as well as shape etc.) and a particular pose (e.g. an orientation, size, location). The appearance and pose of the object(s) and / or background in the edited image generated using the methods described herein have a high accuracy, for example high realism compared to real-life objects and backgrounds with the same appearance and pose. A method of training a system for image processing is also described. In one aspect there is described a method, and a corresponding system, implemented by one or more computers, in particular for image processing. The method includes obtaining at least one neural asset. The at least one neural asset comprises at least one object neural asset representing an object to be depicted in a target image and / or a background neural asset representing a background to be depicted in the target image. Each of the at least one neural asset comprises an appearance token defining the appearance of an object or background and a pose token representing a pose of the respective object or background in the target image. The method further includes generating, using an image generation machine learning model, the target image. The target image is generated by providing, as a conditioning input to the image generation machine learning model, the at least one neural asset. The pose of an object or background may be represented in terms of, for example, a size, orientation and location. By disentangling appearance and pose, objects (or background) can be accurately tracked across time, even when the relative pose changes. This improved accuracy enables complex multi-object real-world scenes to be processed (e.g. edited or generated). In another aspect there is described a method, and a corresponding system, implemented by one or more computers, in particular for training an image processing system. The image processing system may comprise an image generation machine learning model. The method includes obtaining one or more appearance tokens. Each appearance token defines the appearance of an object or background in a training initial image. The method includes obtaining one or more pose tokens. Each pose token defines the pose of the same object or background in a training target image. The method includes combining each respective appearance token and pose token of the one or more appearance tokens and one or more pose tokens to form one or more neural assets. The method includes generating, using the image generation machine learning model, a predicted target image by providing, as a conditioning input to the image generation machine learning model, the one or more neural assets. The method includes training the image generation machine learning model based on an image generation objective that depends on the training target image and the predicted target image, the training including updating learnable parameters of the image generation machine learning model. The pose of an object or background may be represented in terms of, for example, a size, orientation and location. By disentangling appearance and pose, objects (or background) can be accurately tracked across time, even when the relative pose changes. This improved accuracy enables complex multi-object real-world scenes to be processed (e.g. edited or generated). The neural assets may alternatively be referred to as object neural assets in relation to neural assets for objects in a scene, or background neural assets for neural assets that represent a background of a scene. The method may be used to train the image generation machine learning model for editing an image. For example, using the techniques described herein the image generation machine learning model may be trained to edit a given initial image to arrive at an image consistent with the training target image. BRIEF DESCRIPTION OF FIGURES Figure 1 shows an example method of image processing; Figure 2 illustrates examples of 3D-aware editing; Figure 3(a) is an illustration of an example object pose representation; Figure 3(b) is an illustration of an example object pose representation; Figure 3(c) is an illustration of an example object pose representation; Figure 4(a) is a schematic illustration of an example image generation machine learning model which may be used in an example method of image processing; Figure 4(b) is a schematic illustration of an example image generation machine learning model which may be used in an example method of image processing; Figure 4(c) is a schematic illustration of an example image generation machine learning model which may be used in an example method of image processing; Figure 5 illustrates examples of object translation and rotation; Figure 6 illustrates examples of compositional generation; Figure 7 illustrates examples of transfer of backgrounds between scenes; Figure 8 shows an example method for training an image processing system; Figure 9(a) illustrates an example process of obtaining an appearance token and obtaining a pose token which may be performed during an example method for training an image processing system; Figure 9(b) shows an example process of generating a predicted target image which may be performed during an example method for training an image processing system; Figure 10 shows example results demonstrating control of the 3D pose of a single object; Figure 11 shows example results where multiple objects are manipulated; Figure 12 shows a comparison with training on a single frame; Figure 13 shows an example device comprising a processor that is configured to perform an example method of image processing and / or an example method for training an image processing system. DETAILED DESCRIPTION Figure 1 shows an example method of image processing. The method is a computer- implemented method. In S11, the method includes obtaining at least one neural asset. The at least one neural asset comprises at least one object neural asset representing an object to be depicted in a target image and / or a background neural asset representing a background to be depicted in the target image. Each of the at least one neural assets comprise an appearance token defining the appearance of an object or background and a pose token representing a pose of the respective object or background in the target image. The neural assets may alternatively be referred to as object neural assets in relation to neural assets for objects in a scene, or background neural assets for neural assets that represent a background of a scene. The neural assets may be thought of as analogous to 3D assets in the field of computer graphics software.3D assets, or 3D object models, are basic components of a 3D scene in computer graphics software. An example workflow may include selecting N 3D assets from an asset library and placing them into a scene. Formally, one can define a 3D , where Aiis a set of descriptors defining the asset’s appearance and Pidescribes its comprise a set of 3D vertices, faces (in the asset’s canonical frame of reference / pose), and / or UV maps (as will be appreciated by those skilled in the art, U and V refer to the axes of a 2D texture coordinate system). Picontains pose-related information. For example, Pimay comprise information about the 3D asset’s placement in a 3D scene and / or any deviations or deformations from its canonical pose. For example, an object pose token may describe e.g., rigid transformation and scaling from an object’s canonical pose. For example, an object pose token may describe e.g., an absolute pose of an object in its canonical space. A background pose token may describe e.g., a relative camera pose embedding for example. A neural asset may comprise a tuple where is a flattened sequence of K D-dimensional vector and is a D’-dimensional embedding of the neural asset’s pose in a scene. In other words, a may be fully described by embedding vectors, factorized into appearance and pose. As described in more detail below, the embedding vectors may be learnable. This factorization into appearance and pose facilitates independent control over appearance and pose of a neural asset. The appearance token(s) of the object(s) and / or background may be associated with (e.g. derived from) one or more initial images. For example, a first appearance token may be associated with a first object in a first initial image, a second appearance token may be associated with a second object in a second initial image and a third appearance token may be associated with a background in a third initial image. Any of the first, second and third initial images may be the same image or may be different images. In other words, appearance tokens for the object(s) and / or background may be derived from a single initial image or multiple initial images. The method further includes S12, generating, using an image generation machine learning model, the target image. The target image is generated by providing, as a conditioning input to the image generation machine learning model, the at least one neural asset. The method of image processing is capable of processing (e.g. editing, converting) one or more initial image(s) to form a target image. The one or more initial images or the object(s) and / or background that they depict are processed such that at least one of the object(s) and / or background in the at least one initial image is depicted, in an edited form, in the target image. As such, when the object(s) and / or background are derived from one or more initial images, the initial image(s) and target image are related by having the same background and / or at least one same object. The initial image may be referred to as an input image. However, it should be understood that, given that the image generation machine learning model is conditioned on neural assets and not the initial image, the initial images themselves need not be received, e.g. as an input. Figure 2 illustrates some examples of 3D-aware editing which may be performed with neural assets. Given the source image (original image) and object 3D bounding boxes shown on the left of the figure, the figure shows translation, rotation, and rescaling of the object(s). In addition, the figure shows compositional generation, which may be supported by transferring objects or backgrounds across images. In particular, the figure shows replacing of an object(s) and replacing of a background. The target image generated by the image generation machine learning model depicts the object / and or background associated with the neural asset which is provided as a conditioning signal to the image generation machine learning model. The target image is generated such that it depicts the object and / or background’s appearance and the object and / or background’s pose. The target image may be a frame of a video. A video may be generated by generating a sequence of target images, each providing a respective frame of the video. The at least one neural asset may comprise a background neural asset, wherein the background neural asset comprises an appearance token that defines an appearance of a background in the initial image and a pose token that defines a pose of the background in the target image. The pose of the background may be, for example, an orientation, camera angle, zoom level etc. Different poses of the background may represent a translation, rotation or any other movement of a real or simulated image capturing means configured to capture or represent the background. The background may be defined as any region of the initial image which does not contain objects of interest. The pose token associated with the background neural asset may be referred to as a background pose token. The background pose token may be associated with a relative camera pose embedding. The appearance token of the background neural asset may be referred to as a background appearance token. In the context of a background neural asset, the object neural assets may be referred to as foreground neural assets. The at least one neural asset may comprise a plurality of neural assets. As described above, the techniques described herein enable processing and generation of multi-object scenes. The plurality of neural assets may comprise a plurality of object neural assets. Each object neural asset may relate to a different object in the scene (e.g. in the one or more initial images). Additionally or alternatively, the plurality of neural assets may comprise at least one object neural asset and a background neural asset. The techniques described herein allow for an accurate rendering of multiple objects and / or an object against a background. For example, the target image may be generated with an accurate understanding of the appearance and poses of the various object(s) and background and their relative representation in 3D (three-dimensional) space. Due to the accurate understanding, the generated image may be particularly realistic. The image generation machine learning model may be trained to generate a predicted target image given a training neural asset as a conditioning signal, wherein the training neural asset comprises an appearance token derived from a training initial image and a pose token derived from training target image. By training the system on the appearance of an object in a first (initial) frame and the pose of the same object in a different (target) frame, the system is forced to infer underlying 3D structure of an object. Such a paired frame training strategy forces the image generation machine learning model to learn an appearance token that is invariant to object pose and leverage the pose token to synthesize the object in the predicted target image, avoiding the trivial solution of simple pixel-copying. This approach essentially decouples the appearance and pose of objects (or backgrounds), enabling complex multi-object real-world scenes to be processed (e.g. edited or generated). The training initial image and training target image may be selected from a sequence of frames. The sequence of frames may be, for example, sequences of frames of a video. The training initial image may precede the training target image in the sequence. The sequence of frames may be obtained using a sensor, for example a camera. That is, the sequence of frames may comprise video data captured from the real world. The video data may contain, for example, depictions of real-world objects and real-world backgrounds. Video data is particularly useful for training because it represents object-level edits in the real world, i.e. it represents how an object will appear given a particular translation, rotation etc. Video data can accurately demonstrate how objects, scenes, lighting conditions etc. change over time, as well as showing the change in background e.g. due to change in lighting conditions and camera angle or zoom level, for example. Selection may include sampling. The selection or sampling may be, for example, random. The appearance token defining the appearance of an object or background may be derived from an initial image depicting the object or background. The appearance token for each neural asset may be derived from one of multiple initial images. Alternatively, the appearance token for each neural asset may be derived from the same initial image. The method may further comprise generating the appearance token for at least one of the at least one neural asset. Generating the appearance token may comprise processing an initial image using an image encoder to generate, for the object or the background in the initial image, a representation of the appearance of the corresponding object or background. The initial image may contain (e.g. depict) an object or background. One or more 2D bounding boxes bimay be provided for each initial image specifying which object should be extracted. The generated appearance token may define the appearance of the object or background depicted in the initial image. The image encoder may comprise any suitable image encoder, such as a vision transformer (ViT), a contrastive language-image pretrained (CLIP) model, a self-distillation with no labels (DINO) model, a mean absolute error (MAE) based model. The image encoder may be pre-trained, e.g. DINO with self-supervised pre-trained ViT-B. The image encoder may be jointly fine-tuned with the machine learning image generation model in the training process described below. The image encoder used to encode each object and the background may be the same, and / or may share weights. The method may comprise generating the appearance token for each of multiple neural assets. Generating the appearance token for each of multiple neural assets may comprise processing the initial image using the image encoder to generate a representation of the appearance of each object and / or the background in one or more initial image. When multiple object neural assets are received, the appearance token for each object may be derived from the same initial image. Alternatively, the appearance token for each object may be derived from different initial images. The method may comprise processing a first initial image to generate an appearance token for a first neural asset and processing a second initial image to generate an appearance token for a second neural asset. The first and second initial images may be different. The first and second neural assets may relate to a first object and a second object, respectively. The first and second neural assets may relate to an object and a background, respectively. In other implementations, rather than generating the one or more appearance token from one or more initial images, the appearance token may be obtained from a library of appearance tokens. The representation of the appearance of the object or background may comprise a set of embedding vectors output from the image encoder. For example, an object appearance token (i.e. an appearance token associated with an object neural asset) may comprise a set of embedding vectors and a background appearance token (i.e. an appearance token associated with a background neural asset) may comprise a set of embedding vectors. The set of embedding vectors may be arranged such that each embedding vector forms a row of a matrix. The set of embedding vectors may be concatenated to form a single vector. Obtaining the appearance token for a particular object in the initial image may comprise extracting a portion of embedding vectors output from the image encoder. The extraction may be based upon a region of interest associated with the particular object. For example, it may be desired to obtain a set of N object neural asset appearance tokens Aifrom a source image xsrc .xsrcmay be, for example, an image or a frame in a video. The source image xsrcmay be a visual observation. In one implementation, a fully-unsupervised method, such as Slot Attention (Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. NeurIPS, 2020) may be used to decompose an image into a set of object representations. In another implementation, annotations may be used to allow fine-grained specification of objects of interest. For example, a region of interest (such as a 2D bounding box) bimay be provided, obtained or received, for each Neural Asset ai, specifying which object should be extracted from xsrc. The source image may then be filtered based on the specified regions of interests. In one example implementation, the object appearance token Aimay be obtained by , where Hiis an output feature map of a visual encoder, Enc. The image encoder, Enc, is applied to the source image xsrc. The RoIAlign operation (Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In ICCV, 2017) can be used to extract a feature map using the provided region of interest biwhich is flattened to form the object appearance token Ai. In particular, the RoIAlign operation can be used to extract a fixed size feature map using the provided bounding box bi. The RoIAlign operation may be applied to the output of the image encoder Enc and may use the provided bounding box bicorresponding to the object. The output of the RoIAlign operation is then flattened to form the object appearance token Ai. It will be appreciated that the region of interest may take any suitable form such as a point, or any other suitable representation of position such as a textual description. The region of interest may be derived from a plurality of data. The factorization of Aienables extraction of N object appearance tokens from a source image (i.e. initial image) with a single encoder forward pass. That is, the initial image may be filtered using multiple regions of interest. In contrast, previous methods require cropping of each object out of the initial image to extract features separately, and thus require N encoder passes for N objects. The present techniques may be particularly advantageous when jointly fine-tuning the visual encoder, as multiple passes would be computationally infeasible / unaffordable. By jointly training the visual encoder, the present techniques are able to better learn generalizable features of the objects. Obtaining the appearance token for the background in the initial image may comprise receiving an indication of objects of interest in the initial image, each object of interest being associated with a plurality of pixels in the initial image. Obtaining the appearance token for the background in the initial image may further comprise masking the pixels of the initial image associated with each of the objects of interest to generate a masked image. For example, to avoid leakage of foreground object information, all pixels within asset bounding boxes bimay be masked. Obtaining the appearance token for the background in the initial image may further comprise encoding the masked image to generate an encoded representation of the masked image. For example, the masked image is passed through the image encoder Enc. The image encoder Enc may have shared weights with the foreground asset image encoder. A global RoIAlign is then applied, i.e., using the entire image as region of interest, to obtain a background appearance token . When obtaining an appearance token for an object neural asset the output of the image encoder may be filtered based on the region of interest such that only image features associated with the object in the region of interest are extracted. In contrast, with the background appearance token, the entire field of view is considered the region of interest and so all image features of the masked image are extracted. It may be beneficial to encode the background of an image separately to the objects in an image to better facilitate independent control of the background and the objects in target images. By providing a separate background neural asset (i.e. separate to the at least one object neural asset) operations such as swapping the background in a scene, or controlling global properties such as lighting are facilitated. In addition, including background tokens may improve object-level metrics, for example because the model has a reduced requirement to infer background information and / or because typically a large portion of an image comprises background. In some alternative implementations however, the image generation model is not conditioned on any background neural assets, or is conditioned on background appearance tokens but not background pose tokens (e.g. using relative camera pose). The method may use other neural assets associated with other initial images. However, the process of generating the appearance token of the background neural asset associated with the initial image is ambivalent to any objects associated with neural assets derived from other initial images. The pose token of each neural asset may comprise a set of embedding vectors. Each of the pose token and appearance token of a neural asset may comprise a set of embedding vectors. In some implementations, the pose token Piof an object neural asset aimay be the primary interface for controlling the presence and 3D pose of an object in a rendered scene. Obtaining the pose token for an object or background may comprise obtaining a three dimensional bounding shape for the object or background. The three dimensional (3D) bounding shape may be any means of representing the pose of the object in three dimensions. The shape may be referred to instead as a volume or representation. In one example, the 3D bounding shape fully specifies the location, orientation, and size of the object. In one example, the 3D bounding shape is a 3D bounding box. The bounding shape may be a bounding box (e.g. a cuboid). The bounding box may be defined in terms of its corners. Other shapes may be used, for example a sphere and associated orientation, a centre of mass and associated orientation and / or extent. The 3D bounding shape may be received from a user input. For example, user input may be provided using image annotation software or other image editing software. Alternatively, the 3D bounding shape may be determined automatically, e.g. using monocular 3D detection, depth estimation and / or pose tracking. The 3D bounding shape may be selected from a reference image, for example from a frame of a reference video. The reference image may be an image of a particular object but against a different background and / or with a different appearance and / or adjacent different objects compared to the desired target image. The method may include obtaining multiple possible 3D bounding shapes and an associated probability of each possible 3D bounding shape. The method may include selecting, from the multiple possible 3D bounding shapes, a particular 3D bounding shape. The selection may be random or may be based on the associated probabilities. A 3D bounding shape for a background may have a first set of features (e.g. four corners of a box) which bound the entire field of view of the target image. However, additional information contained in the 3D bounding shape (e.g. further corners and / or orientation information) may provide further information relevant to the pose of the background, for example the camera angle, zoom level etc. In some implementations a heuristic strategy may be used to encode the background. In some implementations, a pose token for a background Pbgmay be either a timestep embedding of the video frame (relative to a source frame) or a relative camera pose embedding (if available) for example. The bounding shape may represent a transformed pose of the respective object or background compared to an initial pose of the object or background. The initial pose may be a pose which the object or background has in an initial image. The transformation may comprise, for example, a rotation, translation, reflection or dilation, or any combination thereof. The initial pose of the object or background from an initial image may be determined automatically, for example using monocular 3D detection, depth estimation, pose tracking as described above, or in any other way. The initial pose may be user selected (e.g. user defined, for example using image annotation or image editing software). The method may further comprise obtaining the initial pose of the object or background. The method may further comprise determining the transformed pose of the respective object or background. Determining the transformed pose may be based on standard mathematical principles associated with transformations. The method of image processing described herein can be used to accurately represent an object’s appearance and pose after a transformation. The transformed object in the target image generated using these methods will be rendered with high accuracy (e.g. high realism). Obtaining the pose token for a neural asset may further comprise projecting a two dimensional portion of the bounding shape onto an image plane with a three dimensional depth. The 3D bounding shape may be projected through a multi-layer perceptron (MLP) to produce the pose token. In the example where the 3D bounding shape is a 3D bounding box defined in terms of its corners, four corners spanning the 3D bounding box may be selected and projected onto the image plane to get , with the projected 2D coordinate and the 3D depth . Using a may be beneficial for model for example. The pose representation Pifor a neural asset may be obtained by , where is obtained by concatenating the four selected corners , and is then to using an MLP. Figure 3(a) is an illustration of an example object pose representation where the 3D bounding shape is a 3D bounding box defined in terms of its corners. Figures 3(b) and (c) show two example images. In this example, four corners P0, P1, P2, and P3of a 3D bounding box are projected to the 2D image plane and concatenated to obtain the pose token. The projected four corners form a local coordinate system of the object. In this example, given a 3D bounding box of an object, its four corners are projected to the image space, and their 2D coordinates and depth values concatenated to obtain a 12-D pose vector. The 2D projected points resemble a local coordinate frame for the object, specifying its position, rotation, and scale. On the other hand, the depth is useful for determining the occlusion of objects. Using this representation of object pose may provide more training signal to learn the rotation of objects for example. There are various alternative ways to represent the object pose however, e.g., the coordinate of the 3D box center C with its size and rotation. Alternatively, only three corners are used to define the 3D bounding box for example. Alternatively, the 3D bounding shape may be represented by Fourier coordinate encoding, or any other means of specifying a pose of an object (e.g. the location, orientation and size of an object) within an image to produce the pose token. For example, the Fourier coordinate encoding in Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee, Gligen: Open-set grounded text-to-image generation, In CVPR, 2023; or Xudong Wang, Trevor Darrell, Sai Saketh Rambhatla, Rohit Girdhar, and Ishan Misra, Instancediffusion: Instance- level control for image generation, In CVPR, 2024; may be used. Alternative ways to represent 3D bounding boxes (e.g., concatenation of center, size, and rotation used in 3D object detection, such as in Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom, Pointpillars: Fast encoders for object detection from point clouds, In CVPR, 2019) may be used. As described above, both the appearance Aiand the pose Piof a neural asset aimay be obtained from visual observations (such as an image or a frame in a video). In some implementations, the appearance and pose representations are not encoded from the same observation, e.g., they can be encoded from two separate frames sampled from a video. During training, this allows learning of disentangled and controllable representations, as discussed in relation to the training method described below. In some implementations, rather than generating the one or more appearance tokens from one or more initial images, the appearance token may be obtained from a library of appearance tokens for example. The 3D bounding shape may be received from a user input for example. The appearance token and pose token for a particular neural asset may be concatenated channel-wise. As will be understood by the skilled person, a channel generally refers to a specific feature of an image. In examples where the neural asset is defined as a tuple where the appearance tokens are a flattened sequence of K D-dimensional vector embeddings, each channel may refer to one of the K vector embeddings. Each vector embedding may represent a particular feature of an image (or an object / background thereof). The various features may have been learned by a machine learning model trained to identify, recognise and / or represent features in images, e.g. an image encoder. A feature may be, for example, a distinctive pattern. Each feature may be associated with a respective kernel. When the appearance token and pose token are a background appearance token and a background pose token, the background appearance token may be concatenated channel-wise with the background pose token. Similarly, when the appearance token and pose token are an object appearance token and an object pose token, the object appearance token may be concatenated channel-wise with the object pose token. That is, channel-wise concatenation may be performed for tokens of any object and / or of the background. Channel-wise concatenation of an appearance token and a pose token uniquely binds one pose token with one appearance token to ensure that each object in each image and the background in each image each have a unique representation. Alternatively, positional encoding can be used to learn the association between different objects and their appearance and pose in an image. In some implementations, channel-wise concatenation may be advantageous as learning associations through positional encoding may reduce or break the image generator’s permutation-invariance against the order of input objects and may lead to poorer results. The channel-wise concatenation may be followed by a linear projection. For example, an appearance token Ai and a pose token Pi may be concatenated channel-wise, and then linearly projected to obtain a neural asset representation aias follows: . In an example, the appearance token Ai may be a matrix having size K x D. Pose token Pi, which may be a vector of length D’, may be repeated K times and concatenated with the appearance token Ai. The resulting matrix of size (K x (D + D’)) may then be linearly projected to obtain a neural asset ai. The at least one neural asset may comprise a plurality of neural assets. The method may comprise combining the plurality of neural assets to form a sequence of tokens. It may be the sequence of tokens that is provided as a conditioning input to the image generation machine learning model. Once objects (and / or the background) in an image each have a unique representation (e.g. through channel-wise concatenation or positional encoding), a plurality of neural assets may be combined. The neural assets may be combined using concatenation. The resulting sequence of neural assets (tokens) can be used to condition the image generation machine learning model. The sequence of neural assets may comprise a concatenated sequence of an object neural asset with one or more additional object neural assets and / or a background neural asset. The sequence of neural assets may be used as a drop-in replacement for, by way of example, a sequence of text tokens in a pre-trained text- to-image generation model. The concatenation of multiple neural assets may be performed along the token axis (i.e. the dimension or axis along which individual tokens of a sequence are represented). E.g., instead of text embeddings, an image generation model (e.g. an image generation machine learning model) configured to condition image generation based on text embeddings can instead condition image generation based on the sequence of tokens including multiple neural assets. For example, a set of N neural assets may be encoded into a sequence of tokens that can be appended to or used in place of text embeddings for conditioning an image generation machine learning model. Multiple object neural assets may be concatenated along the token axis to arrive at a token sequence, which can be used as a drop-in replacement for a sequence of text tokens in a text-to-image generation model for example. Similarly, a background pose token may be attached to a background appearance token to generate a background neural asset (for example, a background pose token and a background appearance token are concatenated channel-wise and linearly projected). The object neural assets and the background neural asset may be concatenated along the token dimension and used to condition the image generation machine learning model. The image generation machine learning model may comprise a diffusion model. In some examples, the image generation machine learning model may comprise a latent diffusion model. Generating the target image may comprise initializing the target image or a latent vector representation thereof, by sampling values for the pixels of the target image or for the latent vector representation from a noise distribution. Generating the target image may further comprise, at each of a series of time steps: determining an updated version of the target image or the latent vector representation thereof, by processing the time step and the target image or the latent vector representation thereof, at the time step, using the diffusion model conditioned on the one or more neural assets, to determine a reduced noise version of the target image or of the latent vector representation thereof. Figure 4(a) shows an example image diffusion model 200, which is an example of an image generation machine learning model that may be used in an example method of image processing. The image diffusion model 200 in the example shown is conditioned on object neural assets 203, 205 and 207a and a background token 209, which is a background neural asset. Figure 4(b) shows an example of how during inference, one or more of the object neural assets 203, 205 and 207a can be manipulated to control one or more of the objects 204, 206 and 208 in the generated image 201. In the example shown, the object neural assets 203, 210 and 212 are used to rotate the pose of the third object 208 (blue) and replace the second object 206 by a different object 211 from another image (pink) in the new composed image 202. In more detail, in the example shown in Figure 4(a), the image diffusion model 200 is conditioned on a first object neural asset 203 corresponding to a first object 204 in the generated image 201, which is a car to the left of the generated image 201. The image diffusion model 200 is further conditioned on a second object neural asset 205 corresponding to a second object 206 in the generated image 201, which is a van in the centre of the generated image 201. The image diffusion model 200 is further conditioned on a third object neural asset 207a corresponding to a third object 208 in the generated image 201, which is a car on the right of the generated image 201. The image diffusion model 200 is further conditioned on a background neural asset 209. In the example shown in Figure 4(b), the image diffusion model 200 is again conditioned on the first object neural asset 203, corresponding to the first object 204 in the new composed image 202, which is the car to the left of the new composed image 202. The image diffusion model 200 is further conditioned on a new object neural asset 210, replacing the second object neural asset 205, and corresponding to a new object 211 in the new composed image 202 replacing the second object 206 in the generated image 201. This new object 211 is a car in the centre of the new composed image 202. The image diffusion model 200 is further conditioned on a modified third object neural asset 207b, which is modified from the third object neural asset 207a to rotate the pose, and corresponds to the third object 208 in the new composed image 202, which is the rotated car on the right of the new composed image 202. The image diffusion model 200 is further conditioned on the background neural asset 209. By way of example only, image generation machine learning models (also referred to as image generation ML models, image generation models or generative image models) may include diffusion models such as Imagen which is described in arXiv:2205.11487. An example of implementing a diffusion model in latent variable space is described in arXiv:2112.10752. As another example, a latent diffusion model such as described in Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, High-resolution image synthesis with latent diffusion models, In CVPR, 2022, the entire contents of which are hereby incorporated by reference herein, may be used. Where a moving image is to be generated using a diffusion model, this can be done in various ways. As one example the temporal axis can be treated as an extra spatial dimension. As another example a technique such as that described in arXiv:2402.09470 can be used. In some implementations, the image generation machine learning model is conditioned by the sequence of neural assets by a cross-attention mechanism. A schematic illustration of an example image generation machine learning model which may be used in an example method of image processing is shown in Figure 4(c). The image generation machine learning model is a latent diffusion model. The diffusion model comprises a denoiser ^θ401 which is implemented as a U-Net, which predicts the noise ^ added to the data. In an example, the conditioning token dimension of the example image generation model shown in Figure 4(c) is 1024. In an example, to generate an object neural asset, an image encoder may process an initial image and output a feature map of shape 28x28. A RoIAlign may then be applied to extract a 2x2 small feature map and flatten it, such that the appearance token Aifor the object has dimensionality 4 x D, where D is the number of channels in the feature map. A two-layer MLP (multi- layer perceptron) may be used to transform a 3D bounding box input to a vector of length 1024, corresponding to the pose token Pi in this example. The pose token Pi is repeated 4 times and concatenated with the appearance token Ai, and the resulting representation linearly projected to a vector of length 1024 - the object neural asset. In this example, the conditioning token dimension of the image generation model is 1024, therefore a sequence of object neural assets can be used as a drop-in replacement for a sequence of text tokens. The sequence of neural assets may be mapped to the intermediate layers of the UNet 401 directly via a cross-attention layer. Alternatively, the image generation machine learning model may comprise a domain specific encoder τ, configured to process a tokenised version of a text input. The encoder τ may comprise a transformer for example. The sequence of neural assets may be processed by the encoder τ, which projects a sequence of tokens to an intermediate representation, which is then mapped to the intermediate layers of the UNet 401 via a cross-attention layer. The diffusion model may comprise one or more cross-attention layers. For each of the one or more cross-attention layers, the queries can be obtained from an intermediate representation of the U- Net and the keys and values can be obtained from the sequence of neural assets, for example directly or from an intermediate representation generated by the encoder τ from the sequence of neural assets. In an example, the denoiser may further comprise one or more self-attention layers. In an example, the denoiser comprises two convolutional residual blocks per resolution level. A transformer comprising one or more blocks with alternating layers of (i) self-attention, (ii) a position-wise MLP and (iii) a cross-attention layer is provided at one or more resolution levels between the convolutional blocks. For each of the cross-attention layers, the keys and values can be obtained from the sequence of neural assets. For each of the one or more cross-attention layers, the queries can be obtained from an intermediate representation of the denoiser. For each of the one or more self-attention layers, the queries, keys and values can be obtained from an intermediate representation of the denoiser. Using cross-attention may enable to condition the denoiser on individual neural assets which are vector- based representations, facilitating scene decomposition. For background modeling, all pixels within object boxes may be masked by setting them to a fixed value of 0.5, and features extracted with the same image encoder used to generate the object neural assets for example. The background pose token in this example is obtained by applying a different two-layer MLP on the relative camera pose between the source and the target image. The background pose token is concatenated with the background appearance token, and linearly projected to a vector of length 1024 in this example. In the figure, skip connections are indicated by the dashed lines. Cross attention is indicated by elements 402. Instead of denoising raw pixels, the denoiser is applied to low-dimensional latent code z. Generating a target image may comprise initializing a latent vector representation of the target image by sampling values for the pixels of the target image or for the latent vector representation from a noise distribution. Generating the target image may further comprise, at each of a series of time steps: determining an updated version of the latent vector representation zt-1by processing the time step t and the latent vector representation ztat the time step t, using the diffusion model conditioned on the one or more neural assets, to determine a reduced noise version of the latent vector representation zt-1. For example, the diffusion model output may comprise a prediction of noise, which is used to determine a reduced noise version of the latent vector representation. Each updating iteration may have a corresponding time step t, where an initial time corresponds to an initial noisy latent vector representation, e.g. comprising purely noise, and the final time corresponds to a final latent vector representation generated during inference, e.g. to a supposed fully de-noised latent vector representation. At each updating iteration, the denoising output generated by the diffusion neural network is used to update the current latent representation ztas of the updating iteration t, generating an updated current latent representation zt-1. When the denoising output is a prediction of the noise component, the updated current latent representation zt-1may be determined from the current latent representation zt, the denoising output ^θ, and a noise level for the current updating iteration t. An appropriate diffusion sampler may be used to update the current latent representation, e.g., the DDIM (Denoising Diffusion Implicit Model - Song et al., “Denoising Diffusion Implicit Models”, arXiv:2010.02502v4, October 2022) sampler or another appropriate sampler. After the last updating iteration, the updated current latent representation is output. Optionally, in the last iteration, the updated current latent representation may be generated without use of the sampler. An output image may be generated by processing the output latent vector representation generated at the last updating iteration using a decoder neural network, e.g., one that has been pre-trained in an auto-encoder framework. Alternatively, the image generation model may comprise any other generative image model that accepts a sequence of tokens as a conditioning signal. The image generation model may comprise a convolutional neural network (CNN). The image generation model may be pre-trained, for example the image generation model may be pre-trained to accept a sequence of text prompts as a conditioning input. The at least one object neural asset may be provided in place of, or in addition to, a sequence of text prompts as conditioning input. In some implementations, the learned disentangled representations may enable multi-object scene-level editing as shown in Figures 5 to 7. For example, 3D bounding boxes are encoded to object pose tokens Pi. This may make it possible to move, rescale, and rotate objects by changing the box coordinates. It may also be possible to compose object neural assets aiacross scenes to generate new scenes. In addition, background modeling may support swapping the environment map of the scene. In some implementations, the image generation model has been trained to naturally blend the objects into their new environment at new positions, with realistic lighting effects such as shadowing. Figure 5 illustrates examples of object translation and rotation by manipulating 3D bounding boxes. For example, cars can be translated and rotated in driving scenes. As shown in the examples, objects zoom in and out when moving, and show consistent novel views when rotating. Figure 6 illustrates examples of compositional generation. By composing object neural assets, objects may be removed and segmented, as well as transferred and recomposed between scenes. This figure illustrates examples of compositional generation, where objects are removed, segmented out, and transferred across scenes. In this example, the model handles occlusion and inpaints the scene. Figure 7 illustrates examples of transfer of backgrounds between scenes by replacing the background neural asset. In this example, the objects can adapt to new environments, e.g., the car lights are turned on at night. This figure demonstrates examples of background swapping between scenes. In this example, the image generation model is able to harmonize objects with the new environment. For example, the car lights are turned on and rendered with specular highlight when using a background image from a night scene. In one aspect there is described a method, and a corresponding system, implemented by one or more computers, in particular for training an image processing system. Figure 8 shows an example method for training an image processing system. The image processing system may comprise an image generation machine learning model. In step S21, the method includes obtaining one or more appearance tokens. Each appearance token defines the appearance of an object or background in a training initial image. In step S22, the method includes obtaining one or more pose tokens. Each pose token defines the pose of the same object or background in a training target image. In step S23, the method includes combining each respective appearance token and pose token of the one or more appearance tokens and one or more pose tokens to form one or more neural assets. In step S24, the method includes generating, using the image generation machine learning model, a predicted target image by providing, as a conditioning input to the image generation machine learning model, the one or more neural assets. In S25, the method includes training the image generation machine learning model based on an image generation objective that depends on the training target image and the predicted target image, the training including updating learnable parameters of the image generation machine learning model. The pose of an object or background may be represented in terms of, for example, a size, orientation and location. By disentangling appearance and pose, objects (or background) can be accurately tracked across time, even when the relative pose changes. This improved accuracy enables complex multi-object real-world scenes to be processed (e.g. edited or generated). The neural assets may alternatively be referred to as object neural assets in relation to neural assets for objects in a scene, or background neural assets for neural assets that represent a background of a scene. The method may be used to train the image generation machine learning model for editing an image. For example, using the techniques described herein the image generation machine learning model may be trained to edit a given initial image to arrive at an image consistent with the training target image. An appearance token may be obtained for an object in the training initial image and another appearance token may be obtained for a background in the training initial image, and a pose token may be obtained for the respective object in the training target image and another pose token may be obtained for the respective background in the training target image. That is, the training method may utilise multiple neural assets (belonging to a plurality of objects, an object and a background, or a plurality of objects and a background) for conditioning. The training initial image and the training target image may be selected from a sequence of frames. The sequence of frames may be, for example, sequences of frames of a video. Usually, the object will be present in the initial frame and the target frame. There will be circumstances where the target frame will not contain the object, for example it has gone out of the frame of the image. In these instances, the object will still have a representable appearance and pose, but the appearance and pose will represent the lack of the object in the image frame. The training initial image may precede the training target image in the sequence. By training the system on the appearance of an object in a first (initial) frame and the pose of the same object in a different (target) frame, the system is forced to infer underlying 3D structure of an object. Such a paired frame training strategy forces the image generation machine learning model to learn an appearance token that is invariant to object pose and leverage the pose token to synthesize the object in the predicted target image, avoiding the trivial solution of simple pixel-copying. This approach essentially decouples the appearance and pose of objects (or backgrounds), enabling complex multi-object real-world scenes to be processed (e.g. edited or generated). The method may further comprise selecting (e.g. sampling) the training initial image and the training target image from the sequence of frames. The selection may be random. The correlation between an object in the training initial image and the training target image may be tracked, for example using an object tracking model on an underlying sequence of frames. A pose token (e.g. a background pose token) may be associated with a timestep embedding of a video frame from which it was generated (e.g. relative to a source or reference frame). The appearance token and pose token for a particular neural asset may be combined using channel-wise concatenation. That is, the appearance token and pose token for a particular object or background are combined using channel-wise concatenation. The training initial image and target initial image may comprise images obtained using a sensor. For example, the sensor may be a camera. When the training initial image and target initial image are selected from a sequence of images, the sequence of images may comprise video data. The video data may be captured from the real world and contain, for example, depictions of real-world objects and backgrounds. Video data is particularly useful for training because it represents object- level edits in the real world, i.e. it represents how an object will appear given a particular translation, rotation etc. Video data can accurately demonstrate how objects, scenes, lighting conditions etc. change over time, as well as showing the change in background e.g. due to change in lighting conditions and camera angle or zoom level, for example. The one or more neural assets may comprise a plurality of neural assets. The method may include combining the plurality of neural assets to form a sequence of tokens. The sequence of tokens may be provided as a conditioning input to the image generation machine learning model. The above described methods for obtaining appearance tokens for objects in the initial image (e.g. during inference) can be applied to obtaining appearance tokens for the training initial image. The above described methods for obtaining pose tokens for objects in the target image (e.g. during inference) can be applied to obtaining appearance tokens for the training target image. Obtaining the appearance token for a particular object or background in the training initial image may comprise processing the training initial image using an image encoder to generate, for the particular object or background, a representation of the appearance of the corresponding object or background. Obtaining the pose token for the same particular object or background in the training target image may comprise obtaining a three dimensional bounding shape for the object or background. The image encoder may be jointly fine-tuned with the image generation model. Jointly fine- tuning the image encoder may learn more generalizable appearance tokens in neural assets for example. Figure 9(a) illustrates an example process of extracting neural assets. Pairs of video frames – src image 900 and tgt image 901 - contain objects under different poses. In this example, appearance tokens from a source image 900 are encoded with image enc.904 and RoIAlign 902, and pose tokens are encoded from the objects’ 3D bounding boxes in a target image 901. The appearance and pose tokens are combined to form object neural asset representations. In the figure, a second object neural asset representation 205 is generated from the source image 900 and target image 901. Figure 9(b) shows an image diffusion model 200, which is an example of an image generation model, conditioned on object neural assets 203, 205 and 207, and a separate background neural asset 209, to reconstruct the target image as the training signal. The image diffusion model 200 generates the reconstructed image 201. As described and shown in relation to Figure 3(b) above, during inference, the object Neural Assets 203, 205 and 207 can also then be manipulated to control the objects in a generated image. Instead of denoising raw pixels, in this example a VAE (variational autoencoder) maps images to low-dimensional latent code z, on which the denoiser is applied. Thus in this example, a separate image encoder neural network (e.g., one that has been pre-trained jointly with the decoder described above in relation to Figure 4(c) used to generate images during inference) is also used to encode training images to the latent space. In an alternative implementation, the same image encoder may be used to generate the appearance tokens and to encode training images to the latent space. In the example process shown in relation to Figure 9, training data with 3D annotations is used. Object-level "edits" in 3D space may be used to learn multi-object 3D control capabilities. For example, the training data may comprise video data. In video data, as the camera and the content of the scene moves or changes over time, objects are observed from various view points and thus in various poses and lighting conditions. This signal may be exploited by randomly sampling pairs of frames from video clips, where one frame is taken as the "source" image xsrc900 and the other frame as the "target" image xtgt901. In this example, the appearance token Aiof object neural assets is obtained from the source frame xsrc900 by extracting object features using 2D box annotations. Next, the pose token Pifor each extracted asset is obtained from the target frame xtgt901. In order to do this, the correspondences between objects in both frames is identified. For example, such correspondences can be obtained, for example, by applying an object tracking model on the underlying video. Finally, the image generator 200 is conditioned with the associated appearance and pose representations, and trained to reconstruct the target frame xtgt, e.g., using a denoising loss as described in Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, “High-resolution image synthesis with latent diffusion models”, In CVPR, 2022 for example, the entire contents of which are incorporated by reference herein. The image generator 200 may be conditioned with the associated appearance and pose representations, and trained to reconstruct the target image xtgtfrom a noisy target image generated by adding sampled noise to the target image xtgt. The model components (including the image generation model 200 and the image encoder 904) may be trained jointly using a denoising loss. The denoising loss may measure an error, e.g., a mean-squared error, between (i) a denoising output generated by the image generation model 200 and (ii) a target denoising output. For example, when the denoising output ^θ,is a prediction of the noise added to the data, the target denoising output can be the sampled noise ^. The denoising output generated by the image generation model 200 may comprise a denoising output generated by processing a noisy target image (generated by adding sampled noise to the target image) by the image generation model 200 conditioned on the one or more neural assets. The error between the denoising output and the target denoising output may be used to jointly train the model components, e.g., by determining gradients of the error and then using the gradients to update the parameters of the model components by applying an optimizer. Starting from a pre-trained text-to-image model, the entire model may be fine-tuned end-to-end to accept neural assets tokens instead of text tokens as conditioning signal. Examples Example implementations and comparative examples will now be described. In these examples, images were re-sized to 256x256. First example A first example uses a DINO self-supervised pre-trained ViT-B / 8, as described in Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin, “Emerging properties in self-supervised vision transformers”, In ICCV, 2021, as the visual encoder Enc. In this first example, the visual encoder outputs a feature map of shape 28x28 given a 256x256 image. For each object, a RoIAlign is applied to extract a 2x2 small feature map and flatten it, meaning the appearance token Aifor an object in the first example has a sequence length of K = 4. Since the conditioning token dimension of pre-trained image generation model (described below) used in this first example is 1024, a two-layer MLP (multi-layer perceptron) is used to transform the 3D bounding boxes input to D0= 1024, and linearly project the concatenated appearance and pose token back to 1024. For background modeling, in the first example all pixels within object boxes are masked by setting them to a fixed value of 0.5, and features extracted with the same DINO encoder. The background pose token in the first example is then obtained by applying a different two-layer MLP on the relative camera pose between the source and the target image. This visual encoder model is jointly fine-tuned with the image generation model in the first example. The training used the Adam optimizer with a batch size of 1536 on 256 TPU chips. Training is implemented in JAX using the Flax neural network library. For inference, images are generated by running a DDIM (Denoising diffusion implicit models - Song et al., “Denoising Diffusion Implicit Models”, arXiv:2010.02502v4, October 2022) sampler for 50 steps. In this first example, an image generation model as described in Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, “High-resolution image synthesis with latent diffusion models”, In CVPR, 2022 is used. A diffusion model may be a class of generative models that learns to generate samples by iteratively denoising from a standard Gaussian distribution. In this example, the diffusion model comprises a denoiser ^θ, which is implemented as a U-Net, which predicts the noise ^ added to the data x. Instead of denoising raw pixels, the diffusion model in this example comprises a VAE (variational autoencoder) tokenizer to map images to low-dimensional latent code z, on which the denoiser is applied. In addition, the denoiser is conditioned on text and thus supports text-to-image generation. In this first example, the text embeddings are replaced with neural assets. The model is fine-tuned to support appearance and pose control of 3D objects. Although a specific example is described here, the neural assets may be used with any image generator, for example any image generator that conditions on a sequence of tokens. During training, in this first example all model components are trained jointly using the Adam optimizer with a batch size of 1536 on 256 TPUv5 chips (16GB memory each). A peak learning rate of 5 x10-5is used for the image generator and the visual encoder, and a larger learning rate of 1x10-3is used for remaining layers (MLPs and linear projection layers) in this first example. In this first example, both learning rates were linearly warmed up in the first 1,000 steps and kept constant. A gradient clipping of 1.0 is applied to stabilize training in this first example. In this first example, the model is trained for 200k steps on OBJect and MOVi-E datasets, and 50k steps on Objectron and Waymo Open datasets, since the model in this example overfits more on real-world data with complex backgrounds compared to synthetic datasets. In this example, in order to apply classifier-free guidance (CFG), the appearance and pose token are randomly dropped (setting them as zeros) with a probability of 10%. CFG may improve the performance and also alleviate overfitting in training. As noted above, training is performed in this first example using four datasets with object or camera motion, which span different levels of complexity. The OBJect dataset is described in Oscar Michel, Anand Bhattad, Eli VanderBilt, Ranjay Krishna, Aniruddha Kembhavi, and Tanmay Gupta. Object 3dit: Language-guided 3d-aware image editing, NeurIPS, 2023. The MOVi-E dataset is described in Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator, In CVPR, 2022. The Objectron dataset is described in Adel Ahmadyan, Liangkai Zhang, Artsiom Ablavatski, Jianing Wei, and Matthias Grundmann, Objectron: A large scale dataset of object-centric videos in the wild with pose annotations, In CVPR, 2021. All videos from the bike class are discarded as it contains many blurry frames and inaccurate 3D bounding box annotations. Each video in this dataset comes with object pose tracking throughout the video, and is processed to obtain 3D bounding boxes. Since this dataset does not provide 2D bounding box labels, the eight corners of 3D boxes are projected to the image, and the tight bounding box of projected points is taken as 2D boxes. The Waymo Open dataset is described in Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset, In CVPR, 2020. Only the front view is used and cars that are too small are filtered out. In particular, all cars whose 2D bounding box is smaller than 1% of the image area are removed. Random horizontal flip and random resize crop are applied. The 3D bounding boxes in this dataset only have a heading angle (rotation along the yaw-axis) annotation, and thus the other two rotation angles are treated as 0. The provided 2D boxes and 3D boxes are not aligned, thus 3D boxes are instead projected to get associated 2D boxes. For all datasets, the images are re-sized to 256x256 regardless of the original aspect ratio. During inference, the DDIM sampler is run for 50 steps to generate images. In this first example, the model is found to work well with CFG scale between 1.5 and 4, and thus 2.0 is used in this first example. Figure 10 shows results demonstrating the ability to control the 3D pose of a single object on the OBJect dataset for the first example, and two comparative examples (“Chained” and “3 DIT”). This figure presents results on the unseen object subset. Results are shown on the Translation, Rotation, and Removal tasks. Metrics are computed inside the edited object’s bounding box. Results were averaged over 3 random seeds. Compared to the baseline comparative examples, the first example model does not condition on text (e.g., the category name of the object to edit) and was not pre-trained on multi-view rendering of 3D assets. However, state-of-the-art performance was achieved on all three tasks. This is because the neural assets representation in the first example learns disentangled appearance and pose features, which is able to preserve object identity while changing its placement smoothly. Also, the fine-tuned DINO encoder used in the first example generalized better to unseen objects. Figure 11 shows results on MOVi-E, Objectron, and Waymo Open, where multiple objects are manipulated in each sample. Similar to the single-object case, metrics are computed inside the object bounding boxes. The first example model outperforms baseline comparative examples by a sizeable margin across datasets. The first example model is able to control all objects precisely, preserve their fidelity, and blend them into the background naturally in this example. Since camera pose is encoded, global viewpoint change may also be modelled. The model of the first example is trained on videos, where appearance and pose tokens are extracted from different frames. Figure 12 shows a comparison with training on a single frame. Since the appearance token is extracted by a ViT with positional encoding in the first example, it already contains object position information, which acts as a shortcut for image reconstruction. Therefore, the comparative example model trained on a single frame ignores the input object pose token, resulting in poor controllability. One way to alleviate this may be removing the positional encoding in the image encoder (shown in the figure as the Single NO-PE comparative example), which still underperforms paired frame training. This is because to reconstruct objects with visual features extracted from a different frame, the model may be forced to infer their underlying 3D structure instead of simply copying pixels. In addition, the generator may render realistic lighting effects such as shadows under the new scene configuration. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium. The methods may include methods of inference and methods of training. Figure 12 shows an example device 30 comprising a processor 31 that is configured to perform the methods. Some implementations include a non-transitory computer- readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations that perform any of the methods described herein. In some implementations, the target image generated by the image generation machine learning model may be an image of a real-world environment. The target image may depict one or more real-world objects and / or a real-world background. The initial image may also depict the one or more real-world objects and / or a real-world background. The target image may be an edited version of the initial image. The target image may be edited such that the objects and / or background are edited (for example, have a different pose) but are still realistic and representative of their real-world counterparts. The target image may be generated such that the lighting is realistic, the relational positioning of various objects and the background is realistic. The images described herein (e.g. initial image, target image, training initial image, training target image) may be captured by a camera or other imaging device from a real-world environment. The image generation machine learning model may therefore be trained based on images captured from a real-world environment. In some implementations, the target image generated by the image generation machine learning model may be an image of a simulated environment. The target image may depict one or more simulated objects and / or a simulated background. The initial image may also depict the one or more simulated objects and / or a simulated background. The target image may be an edited version of the initial image. The target image may be edited such that the objects and / or background are edited (for example, have a different pose) but are still realistic and representative of real-world counterparts which the simulated objects and background are intended to emulate. The target image may be generated such that the lighting is realistic, the relational positioning of various objects and the background is realistic. Realistic may be interpreted as consistent with a real-world environment, for example complying with the laws of nature and physics. The images described herein (e.g. initial image, target image, training initial image, training target image) may be simulated. The images may be simulated to emulate real-world objects and backgrounds. The image generation machine learning model may therefore be trained based on images captured from a simulated environment that emulate a real-world environment. In an example, the environment can be an educational environment, e.g., the system can be deployed as part of an education software program that assists a user in learning or practicing one or more corresponding skills. In these examples, the initial image can include an observation of a real- world environment in a first state. The target image can represent a predicted observation of the real- world environment in a second state. The target image can indicate actions for the user to perform or instructions to control equipment in the educational environment. The target image can be used to teach a user regarding the change of state of a real-world environment. The target image may be compared to a reference image, for example a reference image predicted by a user, as a means for testing a user’s predictions regarding the state of a real-world environment. In some further applications the method, or a corresponding system, is used for control of a task in a real-world environment. That is, the initial may relate to the task, e.g. it may comprise a observation of a real-world environment in a first state, and the target image may be used to control e.g. a mechanical system (which may be referred to as a mechanical agent), or a computer system for performing the task. The target image may depict, for example, the pose and appearance of an object when the real-world environment is in a second state different to the first state. The target image may be used to predict or determine an action that the mechanical agent could or should use in order to achieve a desired goal with respect to the object. The action may be, for example, an adjustment. The adjustment may be to change the position of the object and / or to apply some treatment to the object in the position (e.g. pose) it is in during the second state. A mechanical system, hereafter also termed the mechanical agent, may include one or more sensors that capture observations of the environment, e.g., at specified time intervals, as the mechanical agent navigates through the environment or attempts to perform a task in the environment. For example, the observations may include, e.g., one or more of: images, object position data, and sensor data to capture observations as the mechanical agent interacts with the environment, for example sensor data from an image, distance, or position sensor or from an actuator. For example in the case of a robot, the observations may include data characterizing the current state of the robot, e.g., one or more of: joint position, joint velocity, joint force, torque or acceleration, e.g., gravity- compensated torque feedback, and global or relative pose of an item held by the robot. In the case of a robot or other mechanical agent or vehicle the observations may similarly include one or more of the position, linear or angular velocity, force, torque or acceleration, and global or relative pose of one or more parts of the mechanical agent. The observations may be defined in 1, 2 or 3 dimensions, and may be absolute and / or relative observations. The observations may also include, for example, sensed electronic signals such as motor current or a temperature signal; and / or image or video data for example from a camera or a LIDAR sensor, e.g., data from sensors of the mechanical agent or data from sensors that are located separately from the mechanical agent in the environment. The mechanical agent may be associated with a control system that generates control signals for controlling the mechanical agent using the observations generated by the sensors. In particular, the control system may generate control signals that cause the mechanical agent to follow a planned trajectory through the environment by first determining an appropriate action for the mechanical agent to perform, e.g., as part of performing a specified task, e.g., navigating to a particular location, identifying a particular object, moving a particular object to a given location, manipulating a particular object in some way, and so on, and then generating control signals that cause the mechanical agent to perform the action. Such a control system can be deployed on-board the mechanical agent or can be deployed remotely from the mechanical agent and can transmit the control signals to the mechanical agent over a data communication network. The control signals can be control inputs to control the mechanical agent. For example, when the mechanical agent is a robot the control signals can be, e.g., torques for the joints of the robot or higher-level control commands. As another example, when the mechanical agent is an autonomous or semi-autonomous land, air, sea vehicle, the control signals can include actions to control navigation, e.g., steering, and movement of the vehicle, e.g., braking and / or acceleration of the vehicle. For example, the control signals can be, e.g., torques to the control surface or other control elements, e.g., steering control elements of the vehicle, or higher-level control commands. In other words, the control signals can include, for example, position, velocity, or force / torque / acceleration data for one or more joints of a robot or of parts of another mechanical agent (system). In these examples, like the control system, software to implement a method as described herein (“system software”) can be deployed on-board the mechanical agent or can be deployed remotely from the mechanical agent. In some implementations, the mechanical agent control system is an autonomous or semi- autonomous control system, e.g., that autonomously or semi-autonomously controls navigation or other actions of the mechanical agent, e.g. vehicle. Also or instead the mechanical agent control system may have an interface to receive control commands, e.g., from a human operator. In these applications the described system software may be used to provide an additional layer of control, e.g., for safety purposes. For example, the described system software may be used to predict or determine the pose of one or more objects at a future time (e.g. at a future state of the environment). For example the described system software may be used to inhibit control of the mechanical agent in a way that could be dangerous or contrary to one or more rules or preferences that the first computer-implement agent is intended to follow and which are determinable by, at least in part, by assessing the target image, for example with reference to the initial image and / or a reference image. As one example, such rules or preferences (preference scores) may include rules / preferences relating to permitted movement of a vehicle, such as traffic rules, or of a robot, e.g. related to permitted (or forbidden) or preferable rules relating to safe movements or types of task. Such rules / preferences may include rules / preferences relating to decisions to be made to ensure safe behavior of the mechanical agent, e.g., to inhibit damage to the mechanical agent or to a human. As an example, the methods may be used in combination with the control of autonomous vehicles. The initial image may represent an appearance and pose of a vehicle in a real-world environment at a first time. The target image may represent an appearance and pose of the vehicle in the real-world environment at a second time. Using the target image, predictions can be made regarding how to control the autonomous vehicle. For example, the appearance and / or pose of the vehicle in the real-world environment at the second time may be used to determine or select an action to be performed by the vehicle. The action may be the application of an acceleration, deceleration, or turning means. The methods may be used to generate training data for training systems for control of autonomous vehicles. The generated target images may accurately (e.g. realistically) depict objects and their change in pose over time, and so can be used to train systems to accurately predict the movement of objects and the best action to take in such circumstances. In some other implementations, the environment is a real-world environment that includes a manufacturing plant, e.g., a manufacturing plant for manufacturing a product, such as a chemical, biological, or mechanical product, or a food product. As used herein “manufacturing” a product also includes refining a starting material to create a product, or treating a starting material, e.g., to remove pollutants, to generate a cleaned or recycled product. The manufacturing plant may comprise a plurality of manufacturing units such as vessels for chemical or biological substances, or machines for processing solid or other materials. The manufacturing units are configured such that an intermediate version or component of the product is moveable between the manufacturing units during manufacture of the product, e.g., via pipes or mechanical conveyance. In implementations the method or system is used for controlling one or more of the manufacturing units or for controlling movement of the intermediate version or component of the product between the manufacturing units. Thus, in these implementations, the method may comprise obtaining, from one or more sensors, one or more initial images representing a first state of the manufacturing units or of the movement. The sensors may comprise any type of sensor monitoring the manufacturing units or the movement, e.g., sensors configured to sense mechanical movement or force, pressure, temperature; electrical conditions such as current, voltage, frequency, impedance; quantity, level, flow / movement rate or flow / movement path of one or more materials; physical or chemical conditions, e.g., a physical state, shape or configuration or a chemical state such as pH; configurations of the units such as the mechanical configuration of a unit, or valve configurations; image or video sensors to capture image or video observations of the manufacturing units or of the movement; or any other appropriate type of sensor The initial image may relate to an action that controls operation of one or more of the manufacturing units or that controls the movement. The target image may be used to control operation of one or more of the manufacturing units or to control the movement. For example the target image may indicate a required or desired appearance and / or pose of a manufacturing unit and / or product. The target image may be used to determine instructions used to control, e.g., minimize, energy or other resource use, or to control the manufacture to obtain a desired quality or characteristic of the product. For example the actions may include actions that control items of equipment of the plant or actions that change settings that affect the manufacturing units or the movement of the product or intermediates or components thereof, e.g., to adjust or turn on / off items of equipment or manufacturing processes. In some implementations the manufacturing plant has a plant control system to control the manufacturing units or to control the movement. The input may be generated by, e.g. in response to, receiving a control signal from the plant control system and generating an input indicating a request to perform a task based on the control signal. In a similar way to that previously described the plant control system may be autonomous, semi-autonomous, or human-controlled. In a similar way to that previously described the system may implement rules or preferences, e.g., to control or limit energy or other resource allocation, or to ensure a target quality or characteristic of the product, or to constrain operation of the plant, e.g., of the manufacturing units, within safe bounds. In some implementations the environment is the real-world environment of a service facility comprising a plurality of items of equipment, e.g. items of electrical equipment, e.g. electrical components, such as a server farm or data center, for example a telecommunications data center, or a computer data center for storing or processing data, or any service facility. The service facility may also include ancillary control equipment that controls an operating environment of the items of equipment, for example environmental control equipment such as temperature control e.g. cooling equipment, or air flow control or air conditioning equipment. A representation of the state of the environment may be derived from observations made by any sensors sensing a state of a physical environment of the facility or observations made by any sensors sensing a state of one or more of items of equipment or one or more items of ancillary control equipment. These include sensors configured to sense electrical conditions such as current, voltage, power or energy; a temperature of the facility; fluid flow, temperature or pressure within the facility or within a cooling system of the facility; or a physical facility configuration such as whether or not a vent is open. Such conditions and characteristics may be represented in the form of images. The initial image and target image can be used to predict or determine the operation of the facility at different times, e.g., to adjust the operation one or more items of equipment (e.g. electrical components) to control, e.g. minimize, use of a resource, such as a task to control use of electrical power or water. For example, a user can ask which components to turn on to decrease use of the resource, or whether it is safe to turn on or off a given component. The target image can depict the components, items of equipment and / or resources at a particular point in time and subsequently be used to determine instructions regarding operation of the items of equipment based on the generated response, e.g., by turning on or off one or more components. In some implementations, multi-object 3D pose control in image diffusion models is provided. For example, instead of conditioning on a sequence of text tokens, a set of per object representations, object neural assets, are used to control the 3D pose of individual objects in a scene. In some implementations, pre-trained diffusion models are used. In some implementations training is performed on real-world videos to achieve multi-object 3D edits. In some implementations, a self- supervised visual encoder is fine-tuned and connected with a large-scale pre-trained diffusion model, which may scale up to complex real-world data. In some implementations, object neural assets may be obtained by pooling visual representations of objects from a reference image, such as a frame in a video, and may be trained to reconstruct the respective objects in a different image, e.g., a later frame in the video. In some implementations, object visuals are encoded from a reference image while conditioning on object poses from the target frame. This may enable learning disentangled appearance and pose features. Combining visual and 3D pose representations in a sequence-of-tokens format may allow to keep the text-to-image architecture of existing models, with neural assets in place of text tokens. For example, by fine-tuning a pre-trained text-to-image diffusion model with this information, fine-grained 3D pose and placement control of individual objects in a scene may be enabled. In some implementations, object neural assets can be transferred and recomposed across different scenes. Neural assets may enable precise control over the output image for example. In particular, neural assets may enable 3D-aware multi-object control for example. In some implementations, videos of multiple objects may be used as a scalable source of training data for 3D multi-object control. For example, for any two frames sampled from a video, naturally occurring changes in the 3D pose (e.g., 3D bounding boxes) of objects may be treated as training labels for multi-object editing. For example, object neural assets may comprise per object latent representations with consistent 3D appearance but variable 3D pose. In some implementations, object neural assets may be trained by extracting their visual appearances from one frame in a video and reconstructing their appearances in a different frame in the video conditioned on the corresponding 3D bounding boxes. This may support learning consistent 3D appearance disentangled from 3D pose. For example, by training on paired video frames, fine-grained 3D control of individual objects may be enabled. In some implementations, any number of neural assets may be tokenised and this sequence fed in to a fine-tuned conditional image generator for precise, multi-object, 3D control. The neural asset formulation may represent objects with disentangled appearance and pose features for example. The method may be applicable to both synthetic and real-world video datasets, and on 3D-aware single- and multi-object editing tasks for example. In some implementations, neural assets may further support compositional scene generation, such as swapping the background of two scenes and transferring objects across scenes. In some implementations, 3D bounding boxes are leveraged as spatial conditioning, which may enable 3D-aware control such as object rotation and occlusion handling. The neural asset representation may capture both object appearance and 3D pose. For example, a neural asset may comprise an appearance and an object pose representation, trained to reconstruct the object via conditioning a diffusion model. The models may be trained on paired images, meaning disentangled representations may be learned, enabling 3D-aware object editing and compositional generation at inference time. In some implementations, besides the 3D pose of assets, explicit mapping of objects into 3D, such as depth maps or the NeRF (neural radiance field) representation, is not performed. In some implementations, neural assets may comprise vector-based representations of objects and scene elements with disentangled appearance and pose features. For example, an object neural asset may comprise a learnable object centric representation. By connecting with pre-trained image generators, controllable 3D scene generation may be enabled. For example, controlling multiple objects in the 3D space as well as transferring and composing assets across scenes may be enabled, both on synthetic and real-world datasets. In some implementations, datasets that capture other changes in objects, such as deformation (e.g., a walking cat), rigid articulation (e.g., opening of a scissor), and structural decomposition (e.g., tomatoes being cut) may be used. In some implementations, datasets that have 3D bounding box annotations are used. In other implementations, scalable 3D annotation pipelines may be used in place of 3D bounding box annotations for example. Certain novel aspects of the subject matter of this specification are set forth in the claims below, accompanied by further description in Appendix A. This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Thus a system, artificial neural network, or trained artificial neural network as described herein, can be implemented in hardware using electronic circuitry, e.g. in a physical box. Similarly computer code as described herein can be code to emulate such hardware or code for a hardware description language. A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network. In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers. The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers. Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return. Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads. Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework. Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet. The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device. While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination. Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. CLAIMS 1. A computer-implemented method of image processing, the method comprising: obtaining at least one neural asset, the at least one neural asset comprising at least one object neural asset representing an object to be depicted in a target image and / or a background neural asset representing a background to be depicted in the target image, each of the at least one neural asset comprising: an appearance token defining the appearance of an object or background; and a pose token representing a pose of the respective object or background in the target image; generating, using an image generation machine learning model, the target image by providing, as a conditioning input to the image generation machine learning model, the at least one neural asset.
2. The method of claim 1, wherein the at least one neural asset comprises a plurality of neural assets.
3. The method of claim 1 or 2, wherein the image generation machine learning model is trained to generate a predicted target image given a training neural asset as a conditioning signal, wherein the training neural asset comprises an appearance token derived from a training initial image and a pose token derived from training target image.
4. The method of claim 3, wherein the training initial image and training target image are selected from a sequence of frames.
5. The method of any preceding claim, wherein the appearance token defining the appearance of an object or background is derived from an initial image depicting the object or background.
6. The method of any preceding claim, further comprising generating the appearance token for at least one of the at least one neural asset, wherein generating the appearance token comprises:processing an initial image containing an object or background using an image encoder to generate, for the object or the background in the initial image, a representation of the appearance of the corresponding object or background.
7. The method of claim 6, wherein obtaining the appearance token for a particular object in the initial image comprises extracting a portion of embedding vectors output from the image encoder, wherein the extraction is based upon a region of interest associated with the particular object.
8. The method of any claim 6 or 7, wherein obtaining the appearance token for the background in the initial image comprises: receiving an indication of objects of interest in the initial image, each object of interest being associated with a plurality of pixels in the initial image; masking the pixels of the initial image associated with each of the objects of interest to generate a masked image; encoding the masked image to generate an encoded representation of the masked image.
9. The method of any preceding claim, wherein the pose token of each neural asset comprises a set of embedding vectors.
10. The method of any preceding claim, wherein obtaining the pose token for an object or background comprises obtaining a three dimensional bounding shape for the object or background.
11. The method of claim 10, wherein the bounding shape represents a transformed pose of the respective object or background compared to an initial pose of the object or background.
12. The method of claim 10 or 11, wherein obtaining the pose token for a neural asset further comprises projecting a two dimensional portion of the bounding shape onto an image plane with a three dimensional depth.
13. The method of any preceding claim, wherein the appearance token and pose token for a particular neural asset are concatenated channel-wise.
14. The method of any preceding claim, wherein: the at least one neural asset comprises a plurality of neural assets and the method comprises combining the plurality of neural assets to form a sequence of tokens; and it is the sequence of tokens that is provided as a conditioning input to the image generation machine learning model.
15. The method of any preceding claim, wherein the image generation machine learning model comprises a diffusion model, and wherein generating the target image comprises: initializing the target image or a latent vector representation thereof, by sampling values for the pixels of the target image or for the latent vector representation from a noise distribution; and at each of a series of time steps: determining an updated version of the target image or the latent vector representation thereof, by processing the time step and the target image or the latent vector representation thereof, at the time step, using the diffusion model conditioned on the one or more neural assets, to determine a reduced noise version of the target image or of the latent vector representation thereof.
16. A computer-implemented method of training an image processing system comprising an image generation machine learning model, the method comprising: obtaining one or more appearance tokens, each appearance token defining the appearance of an object or background in a training initial image;obtaining one or more pose tokens, each pose token defining the pose of the same object or background in a training target image; combining each respective appearance token and pose token of the one or more appearance tokens and one or more pose tokens to form one or more neural assets; generating, using the image generation machine learning model, a predicted target image by providing, as a conditioning input to the image generation machine learning model, the one or more neural assets; training the image generation machine learning model based on an image generation objective that depends on the training target image and the predicted target image, the training including updating learnable parameters of the image generation machine learning model.
17. The method of claim 16, wherein the training initial image and the training target image are selected from a sequence of frames.
18. The method of claim 16 or 17, wherein the appearance token and pose token for a particular neural asset are combined using channel-wise concatenation.
19. The method of any of claims 16 to 18, wherein the training initial image and target initial image comprise images obtained using a sensor.
20. The method of any of claims 16 to 19, wherein: the one or more neural assets comprises a plurality of neural assets; the method comprises combining the plurality of neural assets to form a sequence of tokens; and it is the sequence of tokens that is provided as a conditioning input to the image generation machine learning model.
21. The method of any of claims 16 to 20, wherein: obtaining the appearance token for a particular object or background in the training initial image comprises processing the training initial image using an image encoder to generate, for the particular object or background, a representation of the appearance of the corresponding object or background; and / or obtaining the pose token for the same particular object or background in the training target image comprises obtaining a three dimensional bounding shape for the object or background.
22. The method of any of claims 16 to 21, wherein the image generation machine learning model comprises a diffusion model, and wherein generating the target image comprises: initializing the target image or a latent vector representation thereof, by sampling values for the pixels of the target image or for the latent vector representation from a noise distribution; and at each of a series of time steps: determining an updated version of the target image or the latent vector representation thereof, by processing the time step and the target image or the latent vector representation thereof, at the time step, using the diffusion model conditioned on the one or more neural assets, to determine a reduced noise version of the target image or of the latent vector representation thereof.
23. A device, comprising a processor that is configured to perform the method of any of claims 1 to 22.
24. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising operations that perform the method of any of claims 1 to 22.