Generating visual data items including multiple elements of specified content

The auto-regressive process in the image generation model addresses the challenge of combining multiple user-specified content elements by iteratively refining data tokens based on textual prompts and concept definition datasets, resulting in realistic visual data items with improved quality and reduced computational costs.

WO2025104308A1PCT designated stage expired Publication Date: 2025-05-22DEEPMIND TECH LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
PCT/EP2024/082593
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-11-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing machine learning models struggle to generate visual data items that incorporate multiple user-specified content elements defined by concept definition datasets, especially without data illustrating the interactions between these elements.

Method used

An auto-regressive process is employed using an image generation model that iteratively refines initial data tokens based on textual prompt items, referencing concept definition datasets to generate visual data items that depict specified content elements, such as participants, backgrounds, and actions.

Benefits of technology

This approach enables the generation of realistic visual data items, including videos, that effectively combine multiple user-defined content elements, avoiding issues like overfitting and attribute mix-ups, and can be done with a lower computational cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082593_22052025_PF_FP_ABST
    Figure EP2024082593_22052025_PF_FP_ABST
Patent Text Reader

Abstract

An image generation model generates a visual data item by auto-regressive process in which a set of initial data tokens is iteratively refined in multiple steps based on corresponding textual prompt items The visual data item is based on an output vector generated in one or more of the steps of the auto-regressive process. The textual prompt items are modified during the auto-regressive process by referencing additional ones of a set of concept definition datasets which define content to be included in the visual data item, so that the visual data item depicts content defined by the referenced concept definition datasets.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATING VISUAL DATA ITEMS INCLUDING MULTIPLE ELEMENTS OF SPECIFIED CONTENTCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Application No. 63 / 600,426, filed on November 17, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND

[0002] This specification relates to data processing using machine learning models.

[0003] Machine learning models receive an input and generate an output, e g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.

[0004] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY

[0005] This specification generally describes a system implemented as computer programs on one or more computers in one or more locations which implements the generation of a visual data item including images of multiple user-specified content items The visual data item may comprise or consist of a video, i.e. a plurality of images (image frames), such as ones which create an impression of continuous motion when viewed successively. Alternatively, the visual data item may comprise or consist of a still image

[0006] One content item may be a participant (a person, an animal or an object) having a user-specified appearance. Alternatively, a given content element may be a user-specified background in the image(s), or a user-specified action performed by a participant. A content item is also referred to here as a “concept”.

[0007] In general terms, a first aspect of the disclosure proposes an auto-regressive (recursive) process in which a set of initial data tokens is iteratively refined in multiple steps by an image generation model based on corresponding textual prompt items, and a visual data item is based on an output vector generated in one or more of the steps of the auto-regressive process (e.g. one of more of the last steps of the process). The textual prompt items are modified during the auto-regressive process by referencing additional one(s) of a set of concept definition datasets which define corresponding content items to be included in the visual data item. Thus, the visual data item depicts corresponding content defined by (all) the referenced concept definition datasets.

[0008] As noted, each concept definition dataset is associated with a corresponding “concept” (a possible item of image content for a visual data item), and may comprise one or more images associated with (and depicting) the concept. For example, if the concept is an object or person (e.g. a specific type of teapot), the corresponding concept definition dataset may be one or more images of the object or person, e.g. shown from respective directions and / or in multiple different configurations or states (e g the specific type of teapot shown: from the side or above or below; and / or with its lid on or off). If the concept is a scene, the corresponding concept definition dataset may be at least one image of that scene.

[0009] The image generation model employs image data tokens (here “data tokens”) which can be generated from corresponding images by an encoder (e.g. a pre-existing encoder) and from which images can be reconstructed using a corresponding decoder (e.g. a pre-existing decoder, which may have been trained with the encoder). In each step of the auto-regressive process, the image generation model may receive an input vector which comprises a certain number of data tokens, encoding one or more images (e.g. m images where m is an integer), and generate an output vector which comprises a certain number of data tokens, encoding one or more images (e.g. q images where q is an integer). The number of images q encoded by the output vector may be less than m (for example, it may only be one image). In the first step of the auto-regressive process, the input vector may be formed from initial data tokens (e.g. empty data tokens, encoding an image with a uniform intensity level). In successive steps, the input vector may comprise (e.g. be composed of) data tokens generated in one or more of the earlier steps. Until the autoregressive process has generated data tokens encoding at least m images, the input vector for steps after the first step may comprise a mixture of data tokens generated in the earlier step(s), and enough of the initial data tokens to complete the input vector for the step (e.g. sufficient data tokens that in total the input vector encodes m images).

[0010] A possible expression of a first aspect of the present disclosure is a computer-implemented method of generating a visual data item comprising one or more images, such as a plurality of images which together form a video. The one or more images each comprise at least one respective value for each of a two-dimensional array of pixels and depict image content defined by a plurality of concept definition datasets. The method employs an image generation model generated based on the concept definition datasets. Initial data tokens are defined (e g. as empty data tokens), and an autoregressive process is performed in a plurality of successive steps. In each step, the image generation model is used to process a corresponding input vector of data tokens to generate a corresponding output vector of data tokens, based on a corresponding textual prompt item for the step. In the first step the input vector comprises the initial data tokens, and in each subsequent step the input vector comprises data tokens from the output vector(s) from one or more of the preceding steps (and in some cases, the input vector may comprise some of the initial data tokens too, e.g. until the auto-regressive process has generated enough data tokens to form a complete input vector). The visual data item is generated based on output vectors generated in one or more of the steps. The textual prompt items each include text tokens associated with the image content items (concepts) represented by (that is, defined by) one or more of the concept definition datasets. The text tokens reference the concept(s) corresponding to the concept definition dataset(s). The concept definition datasets are in a sequence (which is equivalent to saying that the corresponding concepts are in a sequence), and textual prompt items for later said steps include text tokens specifying image content represented by progressively later ones of the concept definition datasets in the sequence. The text tokens associated with a given image content item may be referred at as an “index portion”.

[0011] For example, if the visual data item encodes a video comprising a plurality of image frames, the image frames may be generated from the output vectors generated in one or more of the steps. In an embodiment, the visual data item is generated based on output vectors generated in one or more last step(s) of the auto-regressive process.

[0012] For example, each output vector of the image generation model may encode a single image, so that if the visual data item includes n image frames (where n is a positive integer), these frames may be generated respectively from the output vectors generated in the last n steps of the auto-regressive process. In these steps, the textual prompt item may include text tokens associated with the image content of each of the concept definition datasets which it is desired to include in the visual data item (i.e. index portions for each of the concept definition datasets).

[0013] However, note that a given output vector might encode more than one image frame (i.e. q image frames where q is an integer greater than one). In this case the n image frames of the visual data item may generated based on the output vectors generated in the last n / q steps of the auto-regressive process.

[0014] Alternatively, if the visual data item encodes a single image frame, the image frame may be generated from output vector generated in one of the steps, such as the last step.

[0015] In one case, the set of steps comprises a sequence of non-overlapping step subsets corresponding to respective ones of the sequence of concept definitions datasets, each step subset comprising a plurality of the steps of the auto-regressive process. The textual prompt item may be the same for each step of a given one of the step subsets. The textual prompt input for each of the steps of a given step subset may include text tokens (i.e. an index portion) specifying content defined by a corresponding concept definition dataset in the sequence. Furthermore, it may include text tokens (i.e. an index portion) specifying content associated with each concept definition dataset which is earlier in the sequence than the corresponding definition dataset.

[0016] For example, the respective textual prompt items for the different step subsets may be generated recursively, with the textual prompt item for each step subset (except the first) comprising the textual prompt item for the preceding step subset.

[0017] The concept definition datasets may include at least one background definition (defining a background in a video), at least one subject definition dataset (defining a subject in the video, such as an object or character) and / or at least one action definition dataset (defining an action in a video, e.g. performed by the subject). It has been found that the quality of the video dataset may be affected by the sequence of these concept definition datasets. For that reason, in an implementation the at least one background definition dataset (if any) is earlier in the sequence of datasets that the at least one subject definition dataset (if any), and optionally earlier than the at least one action definition dataset (if any). Similarly, the at least one action definition dataset (if any) may be earlier than the at least one subject definition dataset (if any).

[0018] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.

[0019] The disclosed methods make possible generation of a visual data item, particularly a video, incorporating user-defined content, and in particular more than one element of user-defined content defined by corresponding concept definition datasets, even in the absence of (i.e. without using) any data which illustrates the multiple contentelements interacting with each other. Experimentally, it has been demonstrated that this can be achieved with more realistic visual data items being created than in known methods. For example, it may avoid “overfitting”, in which the visual data item follows the depiction of some of the elements too closely (e.g. if the a concept definition dataset, such as one defining a subject of the visual data item, only shows a certain person standing, only visual data items can be generated in which that person is standing). It may also reduce or avoid a mix-up of attributes as between the content elements (e.g. the visual element depicting one content element specified by a concept definition dataset with the colour of another content element specified by another concept definition dataset).

[0020] Another aspect of the present disclosure, which may be combined with the first aspect, is a method for generating a visual data item performed by an image generation neural network model (“image generation model”) which is obtained using a pre-trained image generation model (e.g. a pre-existing text to video model), such as one pre-trained using concept definition datasets drawn from a large database (e.g. the Internet). The image generation model which performs the method for generating a visual data item may have been obtained by training (“fine-tuning”) an adapted image generation model including the pre-trained image generation model and modified by a modification defined by a plurality of numerical parameters. For example, the pre-trained image generation model may be modified by being supplemented by one or more adapter layers defined by the plurality of numerical parameters. The fine-tuning comprises modifying numerical parameters (e g. defining the adapter layers), and is based on the concept definition datasets which are referenced by the textual prompt items during the generation of the visual data item. For example, the concept definition datasets which are referenced by the textual prompt items during the generation of a visual data item may define content which is a particular example of a more general concept which the pretrained image generation model learnt about during the original training (i.e. a concept defined by a concept definition dataset for the general concept used during the original training of the image generation model).

[0021] For example, the pre-trained image model may have been trained using training data (e.g. from the internet) showing images of many types of a certain class of content (concept) described by natural language (e.g. many types of “tree”), but during the fine-tuning the image generation model learns to generate images of a specific instance of the class (e.g. a specific type of tree) in response to textual prompt items including an index portion for the specific instance of the class having a specific characteristic, e.g. comprising a label associated with the specific instance of the class(i.e. a concept more specific than the concept indicated by the class). The fine-tuning uses a concept definition dataset associated with the specific instance of the class, so that, following the fine-tuning, the fine-tuned image generation model processes a textual prompt item which specifies the image content represented by the corresponding concept definition dataset used in the fine tuning, to generate images including the corresponding specific concept. Thus, during the generation of the visual data item, textual prompt items which have that characteristic cause the image(s) of the visual data item to include (at least one) image of the specific instance of the class.

[0022] In example implementations, the number of numerical parameters which are updated to train (fine-tune) the image generation model based on the concept definition datasets is much lower than the number of trained numerical parameters of the pre-trained image generation model Thus the image generation model can be generated with a computational cost which is much lower than that incurred during the generation of the pre-trained image generation model. Accordingly, multiple users can inexpensively benefit from a computational cost incurred (e g. by a third-party company) to produce the pre-trained image generation model.

[0023] Furthermore, it possible to benefit from generation of the pre-trained image generation model even if it was generated at a time when the concept definition datasets used in the fine-tuning and referenced during the method of generating the visual data item were unavailable. In other words, a rapid and computationally inexpensive process of customizing an image generation model comprising a pre-trained image generation model can be performed when the concept definition datasets become available (e.g. by capturing image(s) they contain), to obtain a trained image generation model operative to generate visual data items including the content defined by the concept definition datasets.

[0024] The present concepts may be expressed as one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the operations of the method explained above.

[0025] Alternatively, the present concepts may be expressed as a computer program product comprising (e.g. one or more computer storage media storing) instructions that when executed by one or more computers cause the one or more computers to perform the operations of the method explained above.

[0026] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Otherfeatures, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE FIGURES

[0027] Exemplary implementations of the present techniques will now be described with reference to the following drawings, in which:

[0028] Fig. 1, which is composed of Figs. 1(a) and 1(b), shows images of two items of content (concepts) to be depicted in a visual data item.

[0029] Fig. 2 depicts schematically the intersection of the manifolds for two concepts in a space of output vectors generated by an image generation model.

[0030] Fig. 3 shows an example of an image generation model.

[0031] Fig. 4 shows a process for generating a visual data item using the image generation model of Fig. 3.

[0032] Fig. 5, which is composed of Figs. 5(a)-(d), shows frames based on output vectors generated using the process of Fig. 4 for the concepts of Fig. 1.

[0033] Fig. 6 is a flow diagram of an example process for generating a visual data item.

[0034] Fig. 7 illustrates schematically a process of training the image generation model ofFig. 3.

[0035] Fig. 8 is a flow diagram of an example process for training an image generation model.

[0036] Fig. 9 shows frames of an example video including the concepts ofFig. 1 and frames of videos generated by two known methods.

[0037] Fig. 10 shows two further items of content (concepts), frames of an example video including the two further items of content, and frames of videos generated by two known methods.

[0038] Fig. 11 shows quantitative results comparing example videos and ones generated by other methods.

[0039] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Computer implemented methods will now be described for visual element generation which are examples of implementations of the present disclosure. These methods may be implemented by one or more computer systems, in one or more locations.

[0041] Text-to-video models are known which create videos based on text supplied by a user which specifies the content of the video. A certain item of content may for example be an object to be depicted in the video, and can be thought of as a “concept”. It would desirable to able to generate visual data items (videos, and / or sets of one or more still images) which contain more than one item of user-specified content, but this can prove challenging. For example, for a given item of user specified content (e.g. “teapot” or “tree”) the set of videos which the text-to-video model can produce which contain this content item can be considered as corresponding to points on a complex manifold in a high dimensional space (the space of all videos). The space may be the space spanned by possible realizations of an output vector (e.g. a set of data tokens) of an image generation model which can be processed to generate a corresponding video. A video which depicts both the content items (e.g. both a “teapot” and a “tree”) is at the intersection of these two manifolds. It can be a challenging task for a text-to-video model to find a point on this intersection, so that from that point the text-to-video model can obtain a video depicting both content items.

[0042] This problem is particularly challenging in the case that it is desired for the text-to-video system to create a video in which the content is specified more precisely than just a common word. For example, rather than the user wants to obtain a visual data item which depicts both a teapot of any kind and a tree of any kind, the user may want to produce a visual data item which depicts both a specific type of teapot (e.g. one with a specific shape and / or color) and a specific type of tree.

[0043] The disclosure proposes that an image generation model (image generation neural network model) generates visual items (typically videos, but in principle still images) by an auto-regressive process, in which a set of initial data tokens is iteratively refined in multiple steps by an image generation model based on a corresponding textual prompt item for each step. The initial data tokens define an input vector for the first step, and each step generates a corresponding output vector which may also be composed of data tokens. For each step except the last, the output vector may be used as the input vector for the next step. A visual data item is based on an output vector generated in one or more of the steps of the auto-regressive process. For example, if the visual data item is a still image, the visual data item may be generatedbased on the output vector for the last of the steps. If the visual data item is a video, frames of the video may be generated based on the output vectors generated in respective ones of the steps (e.g. one or more of the last steps of the process). The textual prompt items are modified during the auto-regressive process by referencing additional one(s) of a set of concept definition datasets which define content to be included in the visual data item.

[0044] That is, a sequence of a plurality of concept definition datasets may be defined, and the textual prompt items for later steps may include textual references to a greater number of the concept definition datasets, specifically textual references to concept definition datasets which are later in the sequence. For example, the corresponding textual prompt items for a first one or more of the steps may be a first textual prompt item which includes textual references to (only) the first one of the sequence of concept definition datasets; the textual prompt item for a second one or more of the steps (e.g. one or more of the steps which immediately follow the first one or more of the steps) may be a second textual prompt item which includes textual references to the first one of the sequence of concept definition datasets and the second one of the sequence; the textual prompt item for a third one or more of the steps (e g. one or more of the steps which immediately follow the second one or more of the steps) may be a third textual prompt item which includes textual references to the first and second ones of the sequence of concept definition datasets and to a third one of the sequence the sequence of concept definition datasets (if any); and so on.

[0045] Each concept definition dataset may include one or more (still or moving) images defining an item of content (e.g. a person, an object or a backgrounds). In some cases, the concept definition database may be indicate the appearance of the item of content as seen from multiple respective directions and / or in multiple configurations (e g. a person or object with at least one element in different translational and / or rotational positions in relation to other elements of the person or object).

[0046] One example of a possible concept definition dataset is a subject definition dataset which defines the appearance of a participant (a “participant subject”) in the video. The participant subject may be a character (e.g. a person, an animal, or a cartoon character), such as a character having a body with relatively movable members (limbs), or may be an object (e.g. a teapot). The subject definition dataset may define the appearance of the participant subject by including one or more (still or moving) images depicting the person. These images may, for example, be real -world images (e.g. of a real person, animal or object) captured by a camera.

[0047] Thus, if multiple subject definition datasets are referenced by the textual prompt items using the process for generating a visual data item, the generated visual data item would (ideally) depict each of the corresponding participant subjects.

[0048] Another example of a possible concept definition dataset is a background definition dataset defining a background of the image, i .e. data defining an environment in which the participant subjects are depicted in the visual data item, and for example in which they perform actions.

[0049] Another example of a possible concept definition dataset is an action definition dataset representative of an action in a video. The action may be an action of an unusual nature (e.g. a specific type of dance, or an action carried out in a specific sport, e.g. a rugby scrum) which an image generation model trained on general images is unlikely to be able to incorporate into a video without specific guidance.

[0050] The textual prompt items may comprise a plurality of text tokens selected from a vocabulary of tokens. The tokens can each represent, e.g., words, wordpieces or characters in a natural. Furthermore, the text tokens may include one or more characters from a non-natural language, e.g. ACSII characters which are not letter characters.

[0051] As noted, the textual prompt items instruct the image generation model to employ content defined by one or more of the concept definition datasets in generating a visual data item, by referencing those concept definition datasets. The textual prompt items may comprise at least one portion (“an index portion”) which is text tokens associated with a corresponding concept definition file, and referencing the corresponding concept (possible content item). The image generation model treats the index portion as an indication of a corresponding item of content (concept) to depict in the visual data element. In particular, the index portion references one of the concept definition datasets. A textual prompt item may also include one or more other portion(s) which indicate a manner in which content defined by the referenced concept definition dataset(s) is employed.

[0052] As noted, it may be desirable for a user to be able to specify the content of the visual data item more precisely than by using a common word (e g. bear / tree). In principle, a user might be able to do this using an index portion which describes the content in detail using natural language (e.g. “a polar bear wearing a top hat” or “a sycamore tree having red leaves and a generally symmetric shape”). Some image generation systems are trained on a large (e.g. public) database of text and images, e.g. ones which will include various types of bears and trees), so they may have information about content which is described in detail. However, the more precisely the user definesthe content verbally, the less likely it is that the database contains data relating to that concept, so the less likely it is that the image generation system will produce a high quality result. Furthermore, at different points in the video, the appearance of the content item may change from one appearance consistent with the description, to another appearance consistent with the description.

[0053] To facilitate inclusion of user-specified content, in an example of the present disclosure at least one index portion of the textual prompt items used in one or more (e.g. all) of the steps may include a label which references a specific concept definition dataset, i.e. a concept definition dataset which relates to a specific content item (i.e. a single person or object or background) rather than to a class of items (teapots / trees). The label may be one or more text tokens which are not natural language, and may comprise text tokens from a non-natural language (e g. text tokens which are not drawn from a natural alphabet, such as the Roman alphabet, and which may be control characters). For example, the label may be in the format “#a” (note that the quotation marks here are not part of the label, either in this example or the examples of labels and index portions given below).

[0054] The presence of the label means that there is less risk of the image generation model trying to satisfy the index portion using knowledge which is based on images in the training data which depict items from the class containing a concept which is in the same class as the specific concept (e g. a class defined by the natural language word(s) which apply to the specific concept). For example, rather than the index portion using simply the word “teapot”, which references a class of objects (all kinds of teapot) having a wide range of appearances, the index portion may include a label which references a specific type of teapot. As described below, the image generation model may have been trained to process an index portion including such a label, such as by the process described below with reference to Figs. 7 and / or 8, in which a pre-existing trained (that is “pre-trained”) image generation model (trained, e.g. by a third party, based on concept definition datasets drawn from a large database, such as the Internet) is fined- tuned based on data relating to the specific concept.

[0055] Additionally, the index portion may include one or more natural language tokens which are text tokens describing, as an item of natural language, a category of content of which the image content is an example. For example, in the case of a subject definition dataset which defines a participant subject who is particular woman, the index portion may be “#a woman”, where “#a” is the label which the image generation training model has been trained to interpret as a label for the subject definition dataset, and the“woman” is a set of natural language tokens defining a category of possible participant subjects (i.e. a category which includes the particular woman). In other examples, an index portion may be “#b man”, where “#b” is a label, “man” is natural language tokens, and the index portion as a whole references a participant subject who is a specific man having an appearance defined by a subject definition dataset (the subject definition dataset corresponding to the label “#b”). In another example, an index portion is “#c teapot”, where “#c” is a label, “teapot” is natural language tokens, and the index portion as a whole references a participant subject which is a particular type of teapot having an appearance defined by a corresponding subject definition dataset (the one corresponding to the label “#c”).

[0056] Fig. 1 shows images of two example items of content associated with respective index portions Each of these images may be one of the images included in the concept definition dataset associated with the index portion. Fig. 1(a) is an image of a teapot of a certain shape (e.g. an image obtained by photographing a real-world teapot). The color image is one of a plurality of images of the brown teapot comprised in a concept definition dataset for the teapot. The concept definition dataset is associated with an index portion “A Bl* teapot” where “Bl*” is a label, and “a teapot” is composed of natural language tokens describing a class (teapots) to which the content belongs in natural language.

[0057] Similarly, Fig. 1(b) is a greyscale depiction of a color image of a certain tree (e.g. an image obtained by photographing a real-world tree). The color image is one of a plurality of images of the tree comprised in a concept definition dataset for the tree. The concept definition dataset is associated with an index portion “A C2@ tree” where “C2@” is a label, and “a tree” is natural language tokens which describe a class (trees) to which the content belongs in natural language.

[0058] Note that in addition to the portions of a textual prompt item which are index portions (i.e. text tokens associated with one of the concept definition datasets), the textual prompt item may include one or more text tokens (e.g. natural language characters) which are not associated with one of the concept definition datasets, but instead reference knowledge the image generation model may be expected to have, e.g. because the image generation model is generated from a pre-trained image generation model as described below, and the pre-trained image generation model may be expected to have this knowledge. For example, the textual prompt item may be “[A] plays tennis with [B]”, where [A] and [B] are index portions (e.g. “#a woman” and “#b man”) which are sets of text tokens associated with (and referencing) two respective subject definitiondatasets, and the words “play tennis” are not associated with one of the concept definition databases, but instead reference knowledge of the pre-training image generation model (in this case, knowledge about what a tennis game looks like).

[0059] Consider two possible content items for the visual data item: “concept 1” defined by a first concept definition dataset, and “concept 2” defined by a second concept definition dataset. Each of these two concepts is associated with a respective manifold, which is the set of output vectors which the image generation model could produce which include the corresponding concept. An example is given in Fig. 2, for the concepts of Fig. 1. The manifold associated with the B 1 * teapot includes a multitude of contexts in which the teapot can be positioned and / or used. Similarly, the manifold associated with the C2@ tree includes many different videos which include the tree.

[0060] Upon receiving an input vector and a textual prompt item referencing the first concept definition dataset, the trained image generation model may generate an output vector which is a point on a first “manifold” associated with concept 1. The first manifold contains variations of images including concept 1, including concept 1 interacting with concept 2. Similarly, upon receiving an input vector and a textual prompt item referencing the second concept definition dataset, the trained image generation model may generate an output vector which is a point on a second “manifold” associated with concept 2, and which contains images including concept 2, including concept 2 interacting with concept 1.

[0061] The (non-linear) intersection of the two manifolds is the set of output vectors representing the joint concept of a Bl* teapot and a C2@ tree. Upon receiving a textual input prompt including portions associated with both the concept definition datasets, a text-to-video model should generate an image at the (non-linear) intersection of the two manifolds. However, this intersection is in a high-dimensional space, and is hard to find.

[0062] As noted, a proposal made by the present disclose is that the textual prompt items differ at different times during the iterative process of generating a visual data item, so as to instruct the image generation model to employ differing amounts of the content defined by the concept definition datasets as the auto-regressive process continues, e.g. increasing amounts of the content. In this way, output generated in steps of an earlier part of the auto-regressive process conditions a later part of the auto-regressive process, in which more of the content is added.

[0063] Specifically, a sequence may be defined of the two concept definition datasets, e.g. the first concept definition dataset and second concept definition dataset inthat order. An autoregressive process may thus be performed using a textual prompt item in initial step(s) which references only concept 1 (thereby specifying the image content represented the first concept definition dataset), and a textual prompt item in later step(s) which references both concept 1 and concept 2 (thereby specifying image content represented by both the first and second concept definition datasets).

[0064] If the textual prompt item for an early portion of the auto-regressive process includes text tokens associated with the first concept definition dataset (i.e. the textual prompt item includes an index portion which references the first concept definition) but not text tokens associated with the second concept definition, then the image generation model generates an output vector encoding images incorporating concept 1 more easily than if the textual prompt item references both concept definition datasets. These condition the image generation model in a later portion of the autoregressive process, for which the textual prompt item includes both text tokens associated with the first concept definition dataset (i.e. the textual prompt item includes an index portion which references the first concept definition) and also text tokens associated with the second concept definition dataset (i.e. the textual prompt item also includes an index portion which references the second concept definition), so that the image generation model can find an intersection point of the two manifolds.

[0065] It has been found that this allows the image generation model to combine the content defined by the referenced concept definition datasets more successfully, e.g. not mixing up characteristics defined by different ones of the concept definition datasets

[0066] Fig. 3 illustrates a possible structure of an image generation model 300 suitable for use in a process according to the present disclosure. It comprises a pre-trained image generation model having a text embedding model 301 and an auto-regressive unit 304. The text embedding model 301 has been trained (e.g. known methods) to receive text tokens 307, and to generate from them text embedding tokens 308. The autoregressive unit 304 has been trained to receive an input vector 303 of data tokens. The data tokens encode pixel-level data about an image, and a plurality of data tokens can be decoded to generate a pixelated image. The input vector 304 defines a number m of images (e.g. data tokens previously generated by the image generation model 300) where m is an integer which is one or more. The auto-regressive unit 304 also receives the text embedding tokens 308 output by the text embedding model 301. The auto-regressive unit is configured to process the data tokens 303 representative of at least one image, conditioned on the text embedding tokens, to generate an output vector 306 of data tokens defining a number of images, denoted q. The updated data tokens 306 are representativeof an image described by the text tokens 307. The image represented by the updated data tokens 303 may be a successive image in a video which follows the image represented by the data tokens 303. Thus, the pre-trained image generation model 301, 304 can be used auto-regressively to generate a video.

[0067] The pre-trained image generation model 301, 304 may for example be a text-to-video model having this form, such as the Phenaki system described in “R. Villegas, et al., “Phenaki: Variable length video generation from open domain textual description”, 2022.

[0068] Another possibility, if the visual data item to be produced is a still image, is that the pre-trained image generation model 301, 304 is a text-to-image model which operates in an auto-regressive manner, e.g. by a series of steps, each conditioned on a textual prompt item. In each of the steps the input vector 303 to the image generation model contains data tokens defining a still image. During each step, the image generation model refines the data tokens of the input vector based on the textual prompt item, to produce an output vector 306 of refined data tokens. Several image generation models of this type are known.

[0069] Alternatively, even if the visual data item is to be a single image, the pretrained model may still be a text-to-video model. The visual data item may be a single frame of a video produced by the trained image generation model 300, e.g. an image encoded by a data tokens produced in the last step of an auto-regressive method described below.

[0070] The data tokens employed in the methods presently disclosed may, for example, be any of the data token formats used in any documents describing known image text-to-video or text-to-image models, i.e. data tokens generated by the pre-trained image generation model and from which image(s) can be reconstructed (e.g. by a decoding process).

[0071] The auto-regressive unit 304 is in the form of one or more transformer layers, e g., as described in A. Vaswani, et al, “Attention is all you need”, 2017, arXiv: 1706.03762. The transformer layers may be arranged in a sequence where the first transformer layer of the sequence receives the input vector of data tokens 303, and each other transformer layer of the sequence receives an input vector which is the output of the preceding transformer layer of the sequence. Each transformer layer multiplies an input to the layer by a corresponding layer matrix to give a first multiplication result, and generates a layer output based on the first multiplication result.

[0072] For example, any one or more of the transformer layer(s) may be a selfattention layer configured to receive an input vector of data token values (e.g. from a corresponding immediately preceding transformer layer), and transform it into an output vector of values based on multiple weight matrices.

[0073] Specifically, the input vector may be transformed into three feature spaces, f, g and h, by multiplying it by respective weight matrices Wf, Wh and Wg, and the features in feature space h are combined by a matrix of weights which are formed from a normalized inner product of the features in feature spaces f and g. The original input vector may be added to a result of this operation (which may first be multiplied by a scalar value), to generate the output vector. Any one or more of the weight matrices Wf, Wh and Wgmay be the layer matrix.

[0074] Alternatively, any one or more of the transformer layer(s) may be a crossattention layer configured to receive an input vector of data token value and the text embedding tokens, and to perform a cross-attention mechanism.

[0075] Specifically, the input vector and text embedding tokens may be transformed into three feature spaces, f, g and h, by multiplying the input vector and text embedding tokens by corresponding ones of the weight matrices Wf, Wh and Wg, and the features in feature space h are combined by a matrix of weights which are formed from a normalized inner product of the features in feature spaces f and g. The original input vector may be added to a result of this operation (which may first be multiplied by a scalar value), to generate the output vector. Any one or more of the weight matrices Wf, Wh and Wgmay be the layer matrix.

[0076] The image generation model 300 differs from the pre-trained image generation model 301, 304 by the addition of one or more adapter layers 305 for corresponding ones of the transformer layers. The adapter layers modify the operation of the auto-regressive unit 304, so that it becomes, in effect, a modified auto-regressive unit 302. The image generation model 300 is described at certain points in this text as an “adapted image generation model” because it was formed from the pre-trained image generation model 301, 304 by modifying the pre-trained image generation model 301, 304 using the adapter layers 305. The adapted image generation model 300 is then fine turned (for example, as described below with reference to Fig. 7 and Fig. 8), to train numerical parameters defining the adapter layers 305 without modifying parameters of the pre-trained image generation model 301, 304.

[0077] Each adapter layer modifies the operation of a corresponding one of the transformer layers, by adding to the first multiplication result of the correspondingtransformer layer a second multiplication result which is a result of multiplying the input to the transformer layer by a corresponding adapter matrix for the adapter layer defined by the plurality of numerical parameters.

[0078] This illustrated in Fig. 3 by the arrows showing that the inputs to ones of the stack of transformer layers of the auto-regressive unit 304 are transmitted to the corresponding ones of the adapter layers 305, where the corresponding second multiplication result is generated using the corresponding adapter matrix, and the corresponding second multiplication result is transmitted back to the corresponding transformer layer, which adds it to the corresponding first multiplication result, to influence the output of the corresponding transformer layer.

[0079] The adapter matrix for each adapter layer may have a rank which is much smaller (e g. at least ten times smaller) than the rank of the corresponding layer matrix The implementation of this possibility is as described by “LoRA: Low rank adaptation of Large Language Models”, E. J. Hu et al, 2021, arXiv:2106.09685.

[0080] Fig. 4 shows a method 400 of generating a visual data item which is an example of the present disclosure. The method 400 may be performed by one or more computer systems in one or more locations.

[0081] The method 400 includes a plurality of steps performed using an image generation model such as the fine-tuned image generation model 300 described above with reference to Fig. 3.

[0082] In a first step, the image generation model 300 receives an input vector 401 of“initial” data tokens. The number ofinitial data tokens may be (at least) the number of tokens which encode a number m of images, where m is an integer which is at least one. The “initial” data tokens of the input vector 401 (indicated by boxes containing a cross) may be set to any value(s), e.g. the same default value for each (“empty tokens”). The image generation model also receives a textual prompt item 402 (“prompt 0”).

[0083] From these two inputs 401, 402, the image generation model 300 generates a corresponding output vector of data tokens 403, by the process described above with reference to Fig. 3. The output vector is composed of data tokens which is (at least) the number of tokens which encode q images, where q may be equal to one or may be an integer greater than one.

[0084] In a second step, the image generation model 300 receives an input vector 404, which includes the output vector 403 from the previous step. If the number of data tokens in the output vector 403 is less than the number of data tokens which encodes m images, i.e. if q is less than m, the initial tokens 405 are supplemented by empty tokens405 to bring the number of data tokens to be equal to the number of data tokens in the input vector 401.

[0085] This process may be repeated for any number of additional steps. In each z-th step, the prompt to the text to video model is a prompt pz. As noted, for the second step, the input vector to the image generation model 300 may as noted include empty tokens, but in each successive z-th step more data tokens have been generated collectively in the preceding steps (i.e. a number of data tokens q(i-l)), so that after a certain number of steps, the input vector for each step can be composed only of data token generated in preceding steps. Generally, each input vector is composed, as soon as this is possible, of data tokens generated from the most recent steps. For example, if q=l, for i m the input vector for the z-th step is composed of the data tokens generated in the preceding m steps.

[0086] Suppose it is desired that the process of Fig. 4 produces a visual data item including a plurality of items of content (concepts). In this case, an order may be defined in those concepts. The steps, labelled by i, may be divided into a number k of step subsets, e g a first step subset i=0, ...ji, a second set subset i=ji+l, ...j2,where ji,j2, ... Jkare ascending integers. In each step of the first step subset, the text prompt item may be the same and may include text tokens referencing (only) the first of the concepts in the sequence (i.e. specify image content represented by the concept definition dataset associated with the first concept). In each step of the second step subset, the text prompt item may be the same and may include text tokens referencing (only) the first and second concepts in the sequence (i.e specify image content represented by the respective concept definition datasets associated with the first and second concepts). In each step of the third step subset (if any), the text prompt item may be the same and may include text tokens referencing (only) the first, second and third concepts in the sequence (i.e. specify image content represented by the respective concept definition datasets associated with the first, second and third concepts). And so on.

[0087] For example, the input text prompt for each step of the first step subset may be “a C2@ teapot”. In this case, the respective output vectors for two steps of the first step subset may be data tokens which, following decoding, are frames such as shown in Fig. 5(a) and 5(b). The black bars in the figures are artefacts caused by a discrepancy between the respective aspect ratios of (i) images in the training data for the fine-tuning, and (ii) images decoded from the data tokens output by the auto-regressive unit 304 (e g. if images in the training data are squares and the images encoded by the data tokens output by the auto-regressive unit 304 are rectangles).

[0088] The input text prompt for each step of the second step subset may be “A Bl* teapot boiling tea under a C2@ tree”. In this case, the respective output vectors for two steps of the second step subset may be data tokens which, following decoding, are frames such as shown in Fig. 5(c) and 5(d). This shows a teapot under the tree with realistic steam rising from the teapot.

[0089] Fig. 6 shows the stages of a method 600 for generating a visual data item. For example, the process 400 is an example of the method 600. The method 600 may be performed by one or more computers in one or more locations.

[0090] In stage 601 a plurality of initial data tokens are defined. For example, they may be defined at random or having a default value.

[0091] In stage 602, a plurality of steps are performed iteratively. In each step, an image generation model to process a corresponding input vector of data tokens to generate a corresponding output vector of data tokens, based on a corresponding textual prompt item for the step. For all steps but the first, the input vector comprises data tokens obtained from the output vector(s) of one or more of the previous steps. The plurality of steps may be composed of a plurality of successive step subsets.

[0092] The textual prompt items each include text tokens specifying the image content represented by one or more of the concept definition datasets The concept definition datasets are in a sequence (equivalent to a sequence of the corresponding concepts), and textual prompt items for later ones of the steps (e.g. later step subsets) include text tokens specifying image content represented by progressively later ones of the concept definition datasets in the sequence (i.e. reference later ones of the corresponding concepts).

[0093] For the steps of each step subset, the text prompt item may be the same. That text prompt item may reference a corresponding one of the concepts. The text prompt item for the first step subset references the first concept of the sequence. For each step subset except the first, the text prompt item may reference a corresponding one of the concepts in the sequence, and the concepts which are (only) earlier in the sequence. For example, for each successive step subset, the corresponding textual prompt item includes the text tokens of the textual prompt item(s) for any preceding step subsets, and additional text tokens. The additional text tokens reference a next one of the sequence of concepts, i.e. text tokens specifying the image content represented by the concept definition dataset for the corresponding next one the sequence of concepts.

[0094] In stage 603, the visual data item is generated from the output vectors generated in step 602, e.g. from the last output vectors generated in step 602. For example,if the output vector of the image generation model comprises data tokens defining q images, and if the visual data item comprises n images, the visual data item may be generated using the output vectors generated in the last n / q steps. Note that in some cases, n may be equal to one (e.g. the visual data item is a single image).

[0095] Fig. 7 shows a method 700 of obtaining an image generation model 300 for use in the process 400 described in Fig. 4. The image generation model is one such as the adapted image generation model 300 of Fig. 3, configured to process an input vector of data tokens and a natural language textual prompt item, and to generate, based on the input vector and the textual prompt item, an output vector of data tokens depicting an image described by the textual prompt item.

[0096] As mentioned above, the adapted image generation model 300 comprises a pre-trained image generation model 301, 304 and adapter layers 305 defined by numerical parameters. That is, the adapter layers 304 are based on respective (normally non-overlapping) subsets of the plurality of numerical parameters.

[0097] The fine tuning is performed using a respective concept definition dataset for each of one or more concepts. Each concept definition dataset defines a respective item of content (the respective concept) for possible inclusion in a visual data item. Each concept definition dataset is a set of (still or moving) images 701 which show that item of content. The number of images 701 per concept definition dataset may be low, e.g. only 1-3 per concept.

[0098] The initial values of the numerical parameters may be set in any way, e g at random or as a default value. The initial values may be very low, e.g. zero, such that the output of the adapted image generation model 300 for a given input to it is substantially equivalent to the output of the pre-trained image generation component 301, 304, because each transformer layer of the adapted image general model initially performs the same function as the corresponding transformer layer of the pre-trained image generation model.

[0099] Each concept definition dataset is also associated with an index portion. This may be in the form of label associated with the concept definition dataset and a natural language tokens describing the category of image content which of which the image content of the associated definition dataset is an example. For example, a certain one of the concepts may be a specific form of teapot. The concept definition dataset is a set of images 701 of this type of teapot. The index portion may be “a Bl* teapot”, where “Bl*” is the label (index portion) associated with the concept definition dataset, and “teapot” is natural language tokens describing the category of image content (i.e. teapots)of which this specific teapot is an example. This concept is learnt using a training textual prompt item which includes the corresponding index portion.[000100] During the fine-training, single ones of the concept definition datasets are chosen in turn. For a selected concept definition dataset, one or more of the images 701 of the selected concept definition dataset are (separately) encoded by a pre-trained encoder 702 to form a corresponding set of data tokens 703 encoding m images. Optionally, in this process a number of the images 701 smaller than m may be used, with multiple copies of the data tokens produced from one or more of the images 701 appearing multiple times in the set of data tokens 703 (e.g. a single one of the training images 701 may be used to generate a corresponding set of data tokens, and the set of data tokens may appear m times in the data tokens 703). The encoder 702 is part of an encoderdecoder pair, which may have been obtained by a known method. It may for example be the C-ViViT encoder of the Phenaki model referenced above.[000101] One or more of the data tokens 703 (e.g. chosen at random) are masked by replacing the data token with an “empty” data token (indicated by an X in Fig. 7) to give a set of masked data tokens 704.[000102] The text tokens of a training textual prompt item including the corresponding index portion for the selected concept definition dataset (e g. the training textual prompt item may be equal to the corresponding index portion) are input to the text embedding model 301.[000103] The set of masked data tokens 704 is received by the auto-regressive unit 304 at a time when the auto-regressive unit 304 is receiving (i.e. is conditioned on) the text embedding tokens generated by the text embedding model 301 from the corresponding training textual prompt item. The auto-regressive unit 304, with its operation modified by the adapter layers 305, generates an output vector 705 of data tokens which could be decoded by the decoder of the encoder-decoder pair to generate q images.[000104] Repeated iterative updates are then made to the plurality of numerical parameters defining the adapter layer 305 to increase the similarity (e.g. according to a similarity measure) between (i) the output vector 705 of the adapted image generation model, and (ii) a plurality of the data tokens 703. For example, the similarity measure may be between (i) the output vector 705 and (ii) a set of the plurality of data tokens 703 which present the last q images of the m images represented by the data tokens 703. Thus, the adapter layer 305 is trained to use the text embedding tokens generated from thecorresponding training textual prompt item to reconstruct the plurality of the data tokens 703 (i.e. to reverse the effect of the masking).[000105] The paragraphs above are in relation to a single one of the concept definition datasets. However, in different ones of the updates, different ones of the concept definition datasets are used, so that gradually the adapted image generation model learns to recognize index portions for all the respective concept definition datasets, and, upon recognizing one, to generate an output vector including data tokens encoding an image incorporating the corresponding content item. The process may end when at least one termination criterion is met. For example, a termination criterion may be that for one or more of the updates, a measure of a magnitude of the updates is below a threshold value, and / or a termination criterion may be that a certain number of updates has been performed.[000106] Since the training textual prompt item may include an index portion which is in the form of a label associated with the image content of the associated concept definition dataset, and natural language text tokens describing a category of image content of which the image content of the associated concept definition dataset is an example, during the training procedure, the image generation model can use the category to assist it in incorporating the content defined by the corresponding concept definition dataset into the image generation model.[000107] For example, if the concept definition dataset is a subject definition dataset defining a character (e g. a specific person), and the category is “person”, the adaptation of the pre-trained image generation model may be such as to modify the aspects of the pre-trained image generation model which are invoked when a textual prompt item it receives specifies a “person”, such that following the modification the image generation model, upon receiving a textual prompt item including an index portion associated with a corresponding index text portion and including the natural language text tokens “person”, is more liable to generate image depicting a person with the appearance defined by the corresponding index portion.[000108] Fig. 8 shows the stages of a method 800 for obtaining an image generation model (e.g. one suitable for use in the methods of Fig. 4 or Fig. 6) based on a plurality of concept definition datasets representing corresponding potential image content. For example, the process 700 is an example of the method 800. The method 800 may be performed by one or more computers in one or more locations.[000109] In stage 801 a pre-trained image generation model is obtained. For example, the image generation model may be composed of the text embedding model 301and auto-regressive unit 304 explained above with reference to Fig.3. The pre-trained image generation model is configured to process an input vector of data tokens and a natural language textual prompt item, and to generate, based on the input vector and the textual prompt item, an output vector of data tokens depicting an image described by the textual prompt item. As described above, the image generation model composed of the text embedding model 301 and auto-regressive unit 304 is capable of doing this auto- regressively.[000110] In stage 802, the pre-trained image generation model is modified to form an adapted image generation model. The modification is defined by a plurality of numerical parameters. In the example of Fig. 3, the modification is performed by adding the adapter layers 305 which are defined by the plurality of numerical parameters.[000111] Stage 803 is an iterative process of updating the plurality of numerical parameters to increase a similarity measure between:(i) an output of the adapted image generation model for an input vector of data tokens associated with one of the concept definition datasets, and a training textual prompt item associated with the one of the concept definition datasets; and(ii) data tokens encoding one or more images defined by the one of the concept definition datasets. For example, in the case of the method 700 between the data tokens 705 and some or all of the data tokens 703 (e.g. enough of the data tokens 703 to define q images).[000112] In the iterative process of stage 803, there may be successive updates for (a) for a given one of the concept definition datasets, different ones of the images of the concept definition dataset, and (b) different ones of the concept definition datasets. This, the numerical parameters are tuned such that adapted image generation model is used to generate, using text tokens associated with any of the multiple concept definition datasets, images containing the corresponding images content.[000113] Experimental results are now presented of the results of the method 400, to generate visual data items which are videos. The generated videos had a resolution of 160x96 pixels. Fig. 9 shows (lower row) the result of generating a video using the method 400 of Fig. 4. The first textual prompt item, for the first step subset, was “A C2@tree”. The second textual prompt item, for the second step subset was “A Bl* teapot boiling under a C2@ tree”.[000114] The upper two rows of Fig. 9 show videos generated for the textual prompt item “A brown teapot boiling tea under a tree” using the Phenaki algorithm, and a DreamBooth-style fine tuning algorithm (N Ruiz et al, “DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation”, 2023). The Phenaki result does not show the tree clearly, and the DreamBooth result is not even a realistic depiction of a teapot.[000115] Fig. 10 shows another example. In this case, the method 700 was carried out using a set of concept definition datasets including concept definition datasets shown in the left of Fig. 10, having corresponding text token index portions “A Bl* blue teddy bear” and “A C2@ forest themed cinema theatre”. The method 400 was then carried out using a first textual prompt item, for the first step subset, which was “A C2@ forest themed cinema theatre”, and a second textual prompt item, for a second step subset, which was “A B 1 * blue teddy bear dancing in a C2@ forest themed cinema theatre”. This gave a video including the frames shown in the bottom row at the right of Fig. 10.[000116] The upper two rows of the right of Fig. 10 show videos generated for the textual prompt item “A blue teddy bear dancing in a forest themed cinema theatre” using the Phenaki algorithm, and a DreamBooth-style fine tuning algorithm (N Ruiz et al, “DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation”, 2023). The Phenaki result does not show a forest-themed cinema theater, and the DreamBooth result does not show a theater at all.[000117] Quantitative results were obtained using videoCLIP scores (H. Xu et al., “VideoCLIP: Contrastive pre-training for zero-shot, video-text understanding”, 2021) and a self-supervised similarity score, DINO (M. Caron et al, “Emerging properties in self-supervised vision transformers, 2021) The former computes the alignment with text, signifying if the concepts and interaction described by the text are present in the generated video. To compute the DINO score, the average over the individual DINO sore was found with respect to each concept in the video. 14 examples were used for subject-subject customization, 32 examples for subject-motion customization, and 32 examples for subject-background customization. To compute the DINO and videoCLIP scores, for example, 8 results were generated, and the generated video with the best videoCLIP score was used to obtain the final result. The results are shown in the left of Fig. 11. Point 901 shows subject-subject results for DreamBooth, and point 902 shows subject-subject results for an example of the present disclosure. Point 903 shows subject-action results for DreamBooth, and point 904 shows subject-action results for an example of the present disclosure. Point 905 shows subject-background results for DreamBooth, and point 906 shows subject-background results for an example of the present disclosure. In all case the DINO score is higher for the example of the present disclosure, and in the subject-action case the videoCLIP score is also higher for the example of the present disclosure.[000118] The right of Fig. 11 shows human evaluation scores, where the “concept” score relates to whether the desired concepts are present in the video, and the “interaction score” relates to whether the concepts are interacting as desired. The reference numerals 901 to 906 have the same meanings as for the left of Fig. 11. In the subject-subject and subject-action cases, the example of the present disclosure obtained a higher concept score, and a higher interaction score.[000119] Some applications of the disclosed method of generating a visual data item are now described. A system which performs the methods disclosed above is referred to here as a visual data item generator (e.g. a video generator).[000120] A first application is to generate a visual data item as sequence of images representing a temporal progression, e.g. the evolution of an environment, e.g. a real world environment. This may be a video which, when played gives the appearance of motion (e.g. at least 25 image frames per second), or may be a series of images representing the state of the environment at times spaced apart by a longer period (e.g. at one second intervals). Since the visual data item incorporates the (e g. user-defined) content defined by the concept definition datasets, the visual data item shows a temporal progression including this content.[000121] For example, the temporal progression may show how one or more objects (defined by corresponding subject definition datasets, e.g. including image(s) of real- world objects captured by a camera) interact in an environment (e.g. defined by a background definition dataset, which may include image(s) of a real-world environment captured by a camera). This makes it possible to visualize this evolution of the environment, for example to understand how a physical process including the object(s) proceeds. Based on the visual data item a decision may be made to insert the objects, in the real-world, into the real-world environment.[000122] For example, if the textual prompt input (e.g. for the final step subset of the auto-regressive process) were “#d object bounces within #e cage”, where “#d ball” and “#e cage# are index portions referencing two objects (a ball and a cage) having an appearance defined by respective subject definition datasets, the visual data item could be a video which depicts whether in this random motion any undesirable configuration of the ball and the cage occurs (e.g. one in which the ball leaves the cage). If it is determined that this is not the case, the scenario of placing the ball inside the cage might be carried out.[000123] In another example, the textual prompt input (e.g. for the final step subset of the auto-regressive process) might be “#f sofa randomly moves in #g room”, where “#fsofa” is an index portion referencing an object (a sofa) defined by a respective subject definition dataset, and “#g room” is an index portion referencing a room in which a sofa may be positions, as defined by a corresponding background definition dataset. The visual data item might be a video (or series of images at spaced apart times) showing the sofa moving to various positions in the room. Based on this, a particular way of positioning the sofa in the room may be selected, and the sofa in real life may be put in that position in the room. This process may be generalized to provide a method of locating any number of objects having an appearance specified by a user (e.g. by capturing one or more images of the object, and defining a subject definition dataset including those images) within an environment specified by a user (e g. by capturing one or more images of the environment, and defining a background definition dataset including those images).[000124] As noted, the initial data tokens may be empty (e g. encode an image which has the same color and intensity for all pixels). Alternatively, the initial data tokens may encode one or a sequence of (e.g. user specified) images, so that the visual data item it generates continues that image or sequence of images. In this way, the visual data item generator may be used to forecast a possible future evolution of a physical system, starting from captured image(s) of that system, where the concept definition datasets describe objects added to the physical system and / or the environment of the system and / or possible actions the objects may perform.[000125] For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.[000126] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly- embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory program carrier for execution by, or to control the operation of, data processing apparatus. Alternatively or in addition, the program instructions can be encoded on an artificially generatedpropagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. The computer storage medium is not, however, a propagated signal.[000127] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include special purpose logic circuitry, e g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.[000128] A computer program (which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.[000129] As used in this specification, an “engine,” or “software engine,” refers to a software implemented input / output system that provides an output that is different from the input. An engine can be an encoded block of functionality, such as a library, a platform, a software development kit (“SDK”), or an object. Each engine can be implemented on any appropriate type of computing device, e.g., servers, mobile phones, tablet computers, notebook computers, music players, e-book readers, laptop or desktop computers, PDAs, smart phones, or other stationary or portable devices, that includes one or more processors and computer readable media. Additionally, two or more of theengines may be implemented on the same computing device, or on different computing devices.[000130] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). For example, the processes and logic flows can be performed by and apparatus can also be implemented as a graphics processing unit (GPU).[000131] Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.[000132] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.[000133] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be usedto provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.[000134] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subj ect matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.[000135] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other [000136] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.[000137] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in theparticular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results, In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.[000138] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

Claims1. A computer-implemented method of generating a visual data item comprising one or more images, the one or more images depicting image content defined by a plurality of concept definition datasets, the method comprising: defining initial data tokens; and performing a plurality of steps of, in each step, using an image generation model to process a corresponding input vector of data tokens to generate a corresponding output vector of data tokens, based on a corresponding textual prompt item for the step, in the first step the input vector comprising the initial data tokens, and in each subsequent step the input vector comprising the output vector from one or more of the preceding steps; and generating the visual data item based on output vectors generated in one or more of the steps; wherein the textual prompt items each include text tokens specifying the image content represented by one or more of the concept definition datasets, the concept definition datasets are in a sequence, and textual prompt items for later said steps include text tokens specifying image content represented by progressively later ones of the concept definition datasets in the sequence.

2. The method of claim 1 in which the visual data item is a video data item encoding a generated video comprising a plurality of image frames, the image frames being generated from the output vectors generated in one or more of the steps.

3. The method of claim 1 or claim 2 in which the corresponding text prompt item for each step except the first comprises the text tokens included in the textual prompt item for the preceding step.

4. The method of any of claims 1 to 3, in which the set of steps comprises a sequence of non-overlapping step subsets corresponding to respective ones of the sequence of concept definition datasets, each step subset comprising one or more of the steps, and the textual prompt item for each step subset including text tokens specifying content associated with the corresponding concept definition dataset, the textual prompt item for each step subset except the first further including text tokens specifying contentassociated with each concept definition dataset which is earlier in the sequence than the corresponding definition dataset.

5. The method of claim 4 when dependent on claim 2 in which the image frames are generated from one or more of the output vectors generated in the last step subset of the sequence of step subsets.

6. The method of any preceding claim in which the text tokens associated with the image content include a label associated with the image content of the associated concept definition dataset.

7. The method of claim 6 in which the text tokens associated with the image content further include natural language text tokens describing a category of content of which the image content is an example.

8. The method of any preceding claim in which the image generation model was generated based on the concept definition datasets.

9. The method of any preceding claim, in which the concept definition datasets include at least one background definition dataset defining a background for an image.

10. The method of any preceding claim, in which the concept definition datasets include at least one subject definition dataset defining the appearance of a participant subj ect.

11. The method of claim 10 in which there are a plurality of subj ect definition datasets defining the appearance of a plurality of corresponding participant subjects.

12. The method of claim 10 or 11 when dependent upon claim 9, in which the at least one background definition dataset is earlier in the sequence of concept definition datasets than the at least one subject definition dataset.

13. The method of any preceding claim when dependent on claim 2, in which the concept definition datasets include at least one action definition dataset representative of an action in a video.

14. The method of claim 13 when dependent upon claim 9, in which the at least one background definition dataset is earlier in the sequence of concept definition datasets than the at least one action definition dataset.

15. A method according to any preceding claim in which one or more of the content definition datasets comprise one or more images captured from the real-world by a camera.

16. A method of obtaining an image generation model based on a plurality of concept definition datasets representing corresponding potential image content, the method comprising: obtaining a pre-trained image generation model configured to process a input vector of data tokens and a natural language textual prompt item, and to generate, based on the input vector and the textual prompt item, an output vector of data tokens depicting an image described by the textual prompt item; modifying the pre-trained image generation model to form an adapted image generation model, the modification being defined by a plurality of numerical parameters; and iteratively updating the plurality of numerical parameters to increase a similarity measure between:(i) an output of the adapted image generation model for an input vector of data tokens and a training textual prompt item associated with one of the concept definition datasets; and(ii) data tokens encoding one or more images defined by the one of the concept definition datasets.

17. The method of claim 16 in which the image generation model comprises a text encoding component trained to process the natural language textual prompt item, and one or more transformer layers configured to generate data tokens autoregressively based on text embedding tokens generated by the text embedding component, the modification being to modify the operation of at least one of the transformer layers using a corresponding adapter layer defined by the plurality of numerical parameters.

18. A method according to claim 17 in which at least one transformer layer is configured to multiply an input to the layer by a layer matrix to give a first multiplication result, and generate an layer output based on the first multiplication result, the adapter layer being configured to add to the first multiplication result a second multiplication result which is a result of multiplying the input to the adapter layer by an adapter matrix defined by the plurality of numerical parameters.

19. A method according to claim 18 in which the adapter matrix has a rank which is at least ten times smaller than the rank of the layer matrix.

20. A method according to any of claims 16 to 19 in which the training textual prompt item includes a label associated with the image content of the associated concept definition dataset and natural language text tokens describing a category of image content of which the image content of the associated concept definition dataset is an example.

21. A method according to any of claims 1 to 15 which is performed using an image generation model generated by a method according to any of claims 16 to 20.

22. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1-21.

23. One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the method of any one of claims 1-21.

Citation Information

Cited By

  • Knitted product image generation method and device based on text adjustment and visual feedback

    CN120219553A

  • Infrared image super-resolution method based on physical guidance visual autoregression

    CN122367744A

  • An infrared image super-resolution method based on physical guidance visual autoregression

    CN122367744B