Apparatus and method for generating an image sequence
The system addresses inefficiencies in generating high-quality image sequences by processing multimodal inputs into latent vectors, enhancing film production efficiency through reduced resource demands and improved quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY SEMICON SOLUTIONS CORP
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Current generative AI technologies struggle to generate high-quality image sequences efficiently, such as videos, due to complex planning and high resource requirements, limiting their practical application in film production.
A system that processes multimodal inputs using encoded tokens, concatenates them into latent vectors, and employs a trained machine-learning model to generate image sequences, leveraging techniques like transformer-based models and video diffusion models to reduce complexity and enhance efficiency.
Enables the generation of high-quality image sequences with reduced computational and resource demands, facilitating more efficient film production by allowing filmmakers to create prefilms that can be refined with minimal human and material resources.
Smart Images

Figure EP2026050787_23072026_PF_FP_ABST
Abstract
Description
[0001] Apparatus and method for generating an image sequence
[0002] Field
[0003] The present disclosure relates to an apparatus for generating an image sequence, a method for generating an image sequence, a method for training a machine-learning model, an apparatus for training a machine-learning model, and a non-transitory machine-readable medium.
[0004] Background
[0005] Video diffusion may refer to an Al (artificial intelligence) algorithm that may be used for artificially generating a video based on an input, such as text, images, or the like. On the other hand, shooting a movie may be very expensive and take a lot of time. It may involve complex planning, trial and error, and then the final execution. Current generative Al technology might be able to generate a movie based on a script, but the quality of such generated movies may be considered too low.
[0006] There may be a demand to provide improved techniques for generating image sequences with Al.
[0007] Summary
[0008] This demand may be satisfied by the independent claims, the drawings, and the following description.
[0009] According to a first aspect, the disclosure provides an apparatus for generating an image sequence based on a multimodal input. The apparatus comprises processing circuitry configured to receive a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated by a trained machine-learning model. The processing circuitry is further configured to concatenate the plurality of encoded tokens to obtain a latent vector for the trained machine-learning model. The processing circuitry is further configured to generate the image sequence using the trained machine-learning model. The machine-learning model receives the latent vector as input.According to a second aspect, the disclosure provides a method fortraining a machine-learning model to generate an image sequence. The method comprises receiving a latent vector that is generated based on concatenating a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated. The method further comprises training the machine-learning model to output the image sequence based on the latent vector.
[0010] According to a third aspect, the disclosure provides a method for generating an image sequence based on a multimodal input. The method comprises receiving a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated by a trained machine-learning model. The method further comprises concatenating the plurality of encoded tokens to obtain a latent vector for the trained machine-learning model. The method further comprises generating the image sequence using the trained machine-learning model. The machine-learning model receives the latent vector as input.
[0011] According to a fourth aspect, the disclosure provides an apparatus for training a machinelearning model to generate an image sequence. The apparatus comprises processing circuitry configured to receive a latent vector that is generated based on concatenating a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated. The processing circuitry is further configured to train the machine-learning model to output the image sequence based on the latent vector.
[0012] According to a fifth aspect, the disclosure provides a non-transitory machine-readable medium comprising machine-readable instruction which, when the instructions are carried out on an apparatus, cause the apparatus to carry out any one of the methods of the second aspect or third aspect.
[0013] Brief description of the Figures
[0014] Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
[0015] Fig. 1 depicts a bock diagram of an apparatus for generating an image sequence according to the present disclosure;
[0016] Fig. 2 depicts a flowchart of a method for training a machine-learning model according to the present disclosure;Fig. 3 depicts block diagram of an apparatus for training a machine-learning model according to the present disclosure;
[0017] Fig. 4 depicts a flowchart of a method for generating an image sequence according to the present disclosure; and
[0018] Fig. 5 depicts a flowchart of a method for generating a prefilm according to the present disclosure.
[0019] Detailed Description
[0020] Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
[0021] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification.
[0022] When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.
[0023] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof,but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.
[0024] Fig. 1 depicts a block diagram of an apparatus 100 for generating an image sequence based on a multimodal input.
[0025] The image sequence may be a series of images presented in a specific order to represent motion, transformation, or change over time. It may be based on an animation, a video, a movie, a film, or the like. Each image in the sequence may capture a distinct moment or phase, which, when viewed in succession, may create the perception of continuity or progression.
[0026] Multimodal input may refer to using or integrating multiple types of data or input modes (in other words: modalities), such as at least two of text, speech, image(s), gesture(s), sensor data, video, audio, and metadata as one input. It may allow artificial intelligence or machinelearning algorithms / models to interpret and process information from different sources simultaneously, thereby enhancing interaction and understanding. For example, a virtual assistant may process spoken commands while also recognizing facial expressions or gestures. Multimodal input may be used in human-computer interaction, robotics, augmented reality, and accessibility technologies.
[0027] The apparatus 100 may be any apparatus that is suitable for processing circuitry according to the present disclosure, such as a computer, server, or the like. The apparatus 100 includes processing circuitry 110. For example, the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a system-on-a-chip (SoC) a neuromorphic processor or a field programmable gate array (FPGA). The processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and / or non-volatile memory. For example, the apparatus 100 may include memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.
[0028] The processing circuitry 110 is configured to receive a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated by a trained machine-learning model.A token may refer to a unit of data used for processing. For example, if the modality is speech or text, the token may represent words, subwords, characters, or the like. Tokenization may refer to the process of breaking down input data (such as speech or text) into smaller units to enable efficient analysis and model training. Tokens may also refer to discrete elements in non-text data, such as pixels in images or feature vectors in structured data. In deep learning models, tokens may be mapped to numerical representations, such as embeddings, for further processing. In more general terms, the token may be based on or correspond to the input modality.
[0029] For the multimodal input described herein, tokenization may include converting the modalities (and tasks) into sequences or sets of discrete tokens, thereby unifying their representation space. When a large multimodal model is used for generating the image sequence, this approach may enable training the model with a single pre-training objective. In such an approach, a task may be formulated as a per-token classification problem, e.g. using crossentropy loss. Thereby, training stability may be enhanced, full parameter sharing may be enabled, and need for task-specific heads, loss function, and loss balancing may be removed. Furthermore, generative tasks may be made more tractable by allowing the model to iteratively predict tokens, either autoregressively or through progressive unmasking. Furthermore, computational complexity may be reduced by compressing dense modalities like images into a sparse sequence of tokens. This may decrease memory and compute requirements. Further information may be found in the scientific publication “4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities” by R. Bachmann et al., published under “arXiv:2406.09406v2” on June 14, 2024.
[0030] A (trained) machine-learning (or machine-learned) model may refer to a data structure and / or set of rules representing a statistical model that the processing circuitry 110 may use at least one machine-learning output, such as at least one token, at least one feature, at least one image sequence, or the like. It should be noted that the present disclosure may be carried out with one machine-learning model that is able to generate the image sequence based on the multimodal input, but the present disclosure is not limited in that regard since different machine-learning models may be used to generate intermediate outputs, as will be apparent from the present disclosure. The data structure and / or set of rules may represent learned knowledge (e.g. based on training performed by a machine-learning algorithm as described below). In machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of training data.The machine-learning model may be trained based on a machine-learning algorithm. The term "machine-learning algorithm" may denote a set of instructions that are used to create, train or use a machine-learning model. For the machine-learning model to determine an output, the machine-learning model may be trained using training data, as commonly known. By training the machine-learning model with a large set of training data and associated training information, the machine-learning model may learn to determine the output. By training the machine-learning model using training information (e.g., multimodal input), the machinelearning model may learns a transformation between training input data and appropriate output data.
[0031] The machine-learning model may be trained using training input data (e.g. data of at least one input modality). For example, the machine-learning model may be trained using a training method called "supervised learning". In supervised learning, the machine-learning model may be trained using a plurality of training samples, wherein each sample may include a plurality of input data values, and a plurality of desired output values, i.e. , each training sample may be associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model may learn which output value to provide based on an input sample that is similar to the samples provided during the training. For example, a training sample may include input data that is based on data of at least one input modality and desired output image sequence. Also, the transformation of a plurality of tokens into a concatenated input may be trained in a similar fashion or with a different approach.
[0032] Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples may lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm or a similarity learning algorithm). Classification algorithms may be used as the desired outputs of the trained machine-learning model may be restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values. Similarity learning algorithms may be similar to classification algorithms but may be based on learning from examples using a similarity function that measures how similar two related objects are.
[0033] Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data may be supplied and an unsupervised learning algorithm may be used to find structure in the input data (e.g. by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data including a plurality of input values into subsets (clusters) so thatinput values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters.
[0034] Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") may be trained to take actions in an environment. Based on the taken actions, a reward may be calculated. Reinforcement learning may be based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
[0035] Furthermore, additional techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and / or the machine-learning algorithm may include a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
[0036] For example, the machine-learning model may be an Artificial Neural Network (ANN). An ANN may refer to a system that is inspired by biological neural networks, such as in retina or in brains. ANNs may include a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There may be three types of nodes: Input nodes that receive input values (e.g., based on at least one modality), hidden nodes that are (only) connected to other nodes, and output nodes that provide output values (e.g., image sequence). Each node may represent an artificial neuron. Each edge may transmit information from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and / or of edges may be adjusted in the learning process. In other words, the training of an ANN may include adjusting the weights of the nodes and / or edges of the ANN, i.e. , to achieve a desired output for a given input.
[0037] Alternatively, the machine-learning model may include a different structure and, e.g., be a support vector machine, a random forest model or a gradient boosting model. Alternatively,the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
[0038] In examples, the machine-learning model may be a combination of any of above examples.
[0039] The circuitry 110 is further configured to concatenate the plurality of encoded tokens to obtain a latent vector for the trained machine-learning model.
[0040] In the following, examples are given for the concatenation of the encoded tokens, but it should be noted that the present disclosure is not limited in that regard.
[0041] For example, tokens from different modalities may be concatenated into a single sequence, e.g., with modality-specific embeddings to indicate their origin. For example, a transformerbased model may be used, such as CLIP (contrastive language image pretraining). Another example for concatenation is using a cross-attention mechanism. In such an approach, instead of direct concatenation, cross-attention layers may be used to allow tokens from one modality to attend to those from another. This may be used in models like Flamingo and BLIP (Bootstrapping Language-Image Pre-training). Another example for concatenation is using shared or modality-specific embeddings. In such an approach, each modality's tokens may be projected into a shared latent space using modality-specific encoders (e.g., CNNs (convolutional neural networks) for images, transformers for text, or the like) before being concatenated. This may ensure that tokens from different domains have compatible feature representations. Another example for concatenation is fusion techniques, such as late fusion (separate processing before combination), early fusion (concatenation at input level), hybrid fusion (a mix of both) determine how tokens are integrated, or the like. Another example of concatenation is positional and modality-specific tokens. In such an approach, to maintain coherence, special tokens (e.g., [IMG], [TEXT]) and positional embeddings may be added to indicate which tokens belong to which modality, helping the model understand and align information effectively.
[0042] A latent vector may refer to a numerical representation of data in a lower-dimensional space (than a dimension of all tokens taken together) that may capture its essential features while removing noise or redundant information. The multimodal inputs like images, text, audio, or the like may thereby be brought into a structured format. Latent vectors may enable efficient data processing, allowing models to learn meaningful relationships and generate new samples from learned distributions.By concatenating multimodal tokens into a latent vector, a unified representation of the tokens may be achieved. Since a latent vector may provide a common feature space where information from different modalities (e.g., text, images, audio) is combined, learning cross-modal relationships may be simplified. Furthermore, dimensionality may be reduced. Instead of processing raw multimodal data (or tokens) separately, concatenating tokens into a latent vector may compress high-dimensional inputs into a more manageable form, thereby enhancing computational efficiency. Furthermore, learning based on a shared latent space may help the multimodal model to generalize across different types of inputs, allowing it to make robust predictions even when one modality is missing or noisy. Furthermore, it may facilitates fusion and alignment. Latent vectors may allow for different modalities to be fused in a structured way, thereby improving tasks like image captioning (aligning visual and textual data), speech-to-text models (aligning audio and text), image sequence generation, or the like. Furthermore, it may enables efficient similarity matching: In models like CLIP, image and text tokens may be mapped to a joint latent space where their similarity can be computed directly, enabling zero-shot learning. However, the present disclosure is not limited to latent vectors, since the tokens may also be concatenated to feature maps, graph representations, attention contexts, sequential representations, hierarchical embeddings, or the like.
[0043] The processing circuitry 110 is further configured to generate the image sequence using the trained machine learning, the machine-learning model receiving the latent vector as input.
[0044] As indicated above, the machine-learning model may be trained to carry out each step, i.e. , the processing (generating) and concatenating of the tokens, as well as the generation of the image sequence. However, according to the present disclosure, different machine-learning models may be applied for each step. For example, generating the multimodal input (i.e., the generation and concatenating of the tokens) may be carried out by a separate (large) multimodal model, whereas the generation of the image sequence may be carried out by a video diffusion model that takes the latent vector as input. Also, the two (or more) models may be unified into one model. It should further be noted that the generation of the tokens of the multiple modalities may be carried out by multiple separate machine-learning models (or encoders) that take the respective data of the respective modality as input and outputs respective tokens, while the multimodal model is trained to concatenate the multiple tokens into the latent vector.
[0045] In other words, in some examples, the processing circuitry 110 is further configured to receive a plurality of inputs of the different modalities. The processing circuitry 110 may be further configured to encode the plurality of inputs based on respective encoders for the respectivemodality to obtain the plurality of encoded tokens. As indicated above, at least one further trained machine-learning model may be used for encoding the respective tokens. Such as (further) trained machine-learning model may be a large multimodal model.
[0046] In some examples, the image sequence is an artificial video that serves as a basis for capturing a real video.
[0047] As indicated above, a video diffusion model may be used for that purpose. A video diffusion model may refer to a generative Al model that synthesizes videos by gradually refining noise into coherent frames through a diffusion process. It may extend diffusion models used in image generation by incorporating temporal consistency, ensuring smooth motion across frames. These models may be trained by learning to reverse a noise-injection process, progressively denoising video sequences to generate realistic outputs.
[0048] In such examples, the apparatus 100 or the circuitry 110 may constitute or be part of a multimodal computer assistant, e.g., for film makers. The assistant may handle a multitude of possible inputs standalone or in combination. Possible inputs are:
[0049] - Text: an entire or partial movie script or partial descriptions of scenes, characters, objects, captions of other input modalities, commands, and the like.
[0050] - Images: photos of artists, photos of pieces of furniture, pictures of landscapes, and the like - Other images, such as sketches
[0051] - Videos of scenes
[0052] - Audio: Voice samples of actors, sound effects, and the like
[0053] - Metadata: length, age restrictions, artistic style, and the like
[0054] Based on the inputs, the assistant may generate a prefilm (as the image sequence). A film director and other members of a film crew may be able to watch the prefilm and make modifications in their respective inputs to obtain a result closer to their imagination. If the filming has already begun, some scenes may be played and filmed and those videos may also be provided as input to the assistant. This prefilm might not become perfect, but it might become a reasonable approximation of the director’s vision. Once the director is satisfied with the result of the prefilm, the main filming may begin (or continued). The film crew may reenact the scenes from the generated prefilm and film this. The basics of a scene may already be provided by the prefilm so that the film crew can focus on the artistic elements of their work. To further simplify this, the assistant may output all technical data for the making of a scene, such as camera angles, movements and specific equipment needed. The director might wantto experiment and request a modification of the prefilm by asking to change an actor or a scene to decide whether it is better to cast another star for a role or whether to film it at a different location.
[0055] Fig. 2 depicts a flowchart of a method 200 for training a machine-learning model to generate an image sequence. The method includes receiving, 210, a latent vector that is generated based on concatenating a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated. The method 200 further includes training, 220, the machine-learning model to output the image sequence based on the latent vector, as has already been discussed above.
[0056] Fig. 3 depicts a block diagram of an apparatus 300 for training a machine-learning model to generate an image sequence. The apparatus 300 includes processing circuitry 310 configured to carry out the method as discussed under reference of Fig. 2. The processing circuitry 310 may be similar or different than the processing circuitry 110 discussed under reference of Fig. 1.
[0057] Fig. 4 depicts a flowchart of a method 400 for generating an image sequence based on a multimodal input. The method includes receiving, 410, a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated by a trained machine-learning model. The method 400 further includes concatenating, 410, the plurality of encoded tokens to obtain a latent vector for the trained machine-learning model. The method 400 further includes generating, 420, the image sequence using the trained machine learning, the machine-learning model receiving the latent vector as input.
[0058] Fig. 5 depicts a flowchart of a method 500 for generating a prefilm (as the image sequence) according to the present disclosure. The method 500 uses data 510 of different modalities which are encoded, 520, into tokens (or features) 530. The tokens 530 are concatenated, 540, into a latent vector (or feature) 550, as discussed herein. A video diffusion model 560 is used for generating a prefilm 570. As discussed above, the prefilm may be imitated by a film crew and actors to generate a final film.
[0059] According to the present disclosure, generative Al may be employed to make movies with less material, human and time ressources than in a conventional approach.
[0060] More details and aspects of the methods described herein are explained in connection with the proposed technique or one or more examples described above (e.g., Figs. 1 to 5). Themethods may include one or more additional optional features corresponding to one or more aspects of the proposed technique or one or more examples described above.
[0061] The following examples pertain to further embodiments of the present disclosure:
[0062] (1) An apparatus for generating an image sequence based on a multimodal input. The apparatus includes processing circuitry configured to receive a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated by a trained machine-learning model. The processing circuitry is further configured to concatenate the plurality of encoded tokens to obtain a latent vector for the trained machine-learning model. The processing circuitry is further configured to generate the image sequence using the trained machine-learning model. The machine-learning model receives the latent vector as input.
[0063] (2) The apparatus of (1), wherein the processing circuitry is further configured to receive a plurality of inputs of the different modalities. In such examples, the processing circuitry is further configured to encode the plurality of inputs based on respective encoders for the respective modality to obtain the plurality of encoded tokens.
[0064] (3) The apparatus (2), wherein the encoded tokens are encoded using a further trained machine-learning model.
[0065] (4) The apparatus of (3), wherein the further trained machine-learning model is a large multimodal model.
[0066] (5) The apparatus of any one of (1) to (4), wherein the input modalities include at least two of text, image, video, audio, and metadata.
[0067] (6) The apparatus of any one of (1) to (5), wherein the image sequence is an artificial video that serves as a basis for capturing a real video.
[0068] (7) The apparatus of any one of (1) to (6), wherein the trained machine-learning model is a video diffusion model.
[0069] (8) A method for training a machine-learning model to generate an image sequence. The method includes receiving a latent vector that is generated based on concatenating a plurality of encoded tokens. Each token represents a modality based on which the image sequenceis to be generated. The method further incudes training the machine-learning model to output the image sequence based on the latent vector.
[0070] (9) The method of (8), wherein the encoded tokens are encoded based on a further trained machine-learning model.
[0071] (10) The method of (9), wherein the further trained machine-learning model is a large multimodal model.
[0072] (11) The method of any one of (8) to (10), wherein the input modalities include at least two of text, image, video, audio, and metadata.
[0073] (12) The method of any one of (8) to (11), wherein the image sequence is an artificial video that serves as a basis for a real video.
[0074] (13) A method for generating an image sequence based on a multimodal input. The method includes receiving a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated by a trained machine-learning model. The method further includes concatenating the plurality of encoded tokens to obtain a latent vector for the trained machine-learning model. The method further includes generating the image sequence using the trained machine-learning model, the machine-learning model receiving the latent vector as input.
[0075] (14) The method of (13), further including receiving a plurality of inputs of the different modalities. The method further includes encoding the plurality of inputs based on respective encoders for the respective modality to obtain the plurality of encoded tokens.
[0076] (15) The method of (14), wherein the encoded tokens are encoded based on a further trained machine-learning model.
[0077] (16) The method of (15), wherein the further trained machine-learning model is a large multimodal model.
[0078] (17) The method of any one of (13) to (16), wherein the input modalities include at least two of text, image, video, audio, and metadata.(18) The method of any one of (13) to (17), wherein the image sequence is an artificial video that serves as a basis for capturing a real video.
[0079] (19) The method of any one of (13) to (18), wherein the trained machine-learning model is a video diffusion model.
[0080] (20) An apparatus for training a machine-learning model to generate an image sequence. The apparatus includes processing circuitry configured to receive a latent vector that is generated based on concatenating a plurality of encoded tokens. Each token represents a modality based on which the image sequence is to be generated. The processing circuitry is further configured to train the machine-learning model to output the image sequence based on the latent vector.
[0081] (21) A non-transitory machine-readable medium comprising machine-readable instruction which, when the instructions are carried out on an apparatus, cause the apparatus to carry out any one of the methods of (8) to (12) or (13) to (19).
[0082] (22) A computer program comprising instructions which, when the computer program is executed on a computer, causes the computer to carry out any one of the methods of (8) to (12) or (13) to (19).
[0083] The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
[0084] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and / or contain machine-executable, processorexecutable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field)programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
[0085] It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, -functions, -processes or -operations.
[0086] If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
[0087] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.
Claims
ClaimsWhat is claimed is:
1. An apparatus for generating an image sequence based on a multimodal input, the apparatus comprising processing circuitry configured to:receive a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated by a trained machine-learning model;concatenate the plurality of encoded tokens to obtain a latent vector for the trained machinelearning model; andgenerate the image sequence using the trained machine-learning model, the machine-learning model receiving the latent vector as input.
2. The apparatus of claim 1 , wherein the processing circuitry is further configured to:receive a plurality of inputs of the different modalities; andencode the plurality of inputs based on respective encoders for the respective modality to obtain the plurality of encoded tokens.
3. The apparatus of claim 2, wherein the encoded tokens are encoded using a further trained machine-learning model.
4. The apparatus of claim 3, wherein the further trained machine-learning model is a large multimodal model.
5. The apparatus of claim 1, wherein the input modalities include at least two of text, image, video, audio, and metadata.
6. The apparatus of claim 1 , wherein the image sequence is an artificial video that serves as a basis for capturing a real video.
7. The apparatus of claim 1, wherein the trained machine-learning model is a video diffusion model.
8. A method for training a machine-learning model to generate an image sequence, the method comprising:receiving a latent vector that is generated based on concatenating a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated; andtraining the machine-learning model to output the image sequence based on the latent vector.
9. The method of claim 8, wherein the encoded tokens are encoded based on a further trained machine-learning model.
10. The method of claim 9, wherein the further trained machine-learning model is a large multimodal model.
11. The method of claim 8, wherein the input modalities include at least two of text, image, video, audio, and metadata.
12. The method of claim 8, wherein the image sequence is an artificial video that serves as a basis for a real video.
13. A method for generating an image sequence based on a multimodal input, the method comprising:receiving a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated by a trained machine-learning model;concatenating the plurality of encoded tokens to obtain a latent vector for the trained machinelearning model; andgenerating the image sequence using the trained machine-learning model, the machinelearning model receiving the latent vector as input.
14. The method of claim 13, further comprising:receiving a plurality of inputs of the different modalities; andencoding the plurality of inputs based on respective encoders for the respective modality to obtain the plurality of encoded tokens.
15. The method of claim 14, wherein the encoded tokens are encoded based on a further trained machine-learning model.
16. The method of claim 15, wherein the further trained machine-learning model is a large multimodal model.
17. The method of claim 13, wherein the input modalities include at least two of text, image, video, audio, and metadata.
18. The method of claim 13, wherein the image sequence is an artificial video that serves as a basis for capturing a real video.
19. An apparatus for training a machine-learning model to generate an image sequence, the apparatus comprising processing circuitry configured to:receive a latent vector that is generated based on concatenating a plurality of encoded tokens, each token representing a modality based on which the image sequence is to be generated; andtrain the machine-learning model to output the image sequence based on the latent vector.
20. A non-transitory machine-readable medium comprising machine-readable instruction which, when the instructions are carried out on an apparatus, cause the apparatus to carry out the method of claim 13.