Method and apparatus for generating composite video data from textual cues
By decomposing text prompts into sub-prompts and combining interpolation and temporal attention regularization techniques, the shortcomings of existing models in generating drastic visual changes are addressed, more coherent video generation is achieved, and the training and verification effects of machine learning models are improved.
Patent Information
- Application Number
- CN202510310990.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-19
AI Technical Summary
Existing text-to-video generation models cannot effectively synthesize visual processes involving drastic visual changes, and tend to generate a single state without temporal transformation.
A large language model is used to decompose text prompts into multiple descriptive sub-prompts, and video data is generated through a video diffusion model. Interpolation and temporal attention regularization techniques are combined to optimize the video generation process.
It achieves fine-grained description and temporal coherence of visual processes, generates more attractive video effects, and enhances the quality of training and validation data for machine learning models.
Smart Images

Figure CN120676111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for generating synthetic video data based on text prompts, in particular for providing video data for training and / or testing and / or verifying and / or validating machine learning models. Background Art
[0002] Generative AI has attracted great interest in recent years. Many efforts have been devoted to text-to-image (T2I) generation, and recent works such as AnimateDiff
[15] and VideoCrafter
[16] have attempted to extend the T2I backbone by adding a temporal module for video generation.
[0003] More specifically, they build on the popular T2I model Stable Diffusion (SD), a diffusion model trained in the latent space of an autoencoder. However, SD's 2D UNet can only process batches of image data. To make the model suitable for video generation, the Text to Video Diffusion model (T2V) introduces a temporal transformer to capture temporal coherence.
[0004] However, current T2V models are unable to synthesize visual processes involving dramatic visual changes and tend to generate a single state without gradual evolution. For example, for a cue like "a boy is getting older," known T2V models tend to only generate the initial state, "a little boy," and therefore fail to show any appropriate temporal transformations as the boy ages. This is attributed to suboptimal language guidance and impaired temporal attention.
[0005] Stable Diffusion (SD) operates not in image space, but in the latent space of an autoencoder. First, the autoencoder E maps a given image x into a spatial latent code z = E(x), and then the decoder D maps z back to the image space. The autoencoder is trained to reconstruct the given image, that is, D(E(x)) ≈ x.
[0006] In the second phase, a diffusion model is trained on this latent space Z. The diffusion model consists of a forward diffusion process and a backward denoising process. Thus, the forward pass is a Markov chain that gradually adds Gaussian noise to the clean data. Therefore, the noise latent quantity can be calculated in closed form. The backward denoising process can be parameterized by another Gaussian distribution.
[0007] During the denoising process, the diffusion model can be conditioned on an additional input vector. In SD, this additional input is typically the text encoding produced by the frozen CLIP text encoder. Given the prompt and its encoding vector, a UNet model is trained to minimize a specific loss function.
[0008] Furthermore, video diffusion models (VDMs), such as VideoCrafter1 and VideoCrafter2, are built on top of the existing T2I model (i.e., stable diffusion), which consists of convolutional layers and spatial transformers. To capture the temporal coherence of video generation, VDM additionally introduces a temporal transformer. Given a desired number of video frames, VDM jointly processes all frames in a batch, i.e., denoising based on the spatial size and feature dimension of each frame. Inheriting from the T2I model, the text cues are first encoded by a frozen text encoder such as the CLIP text encoder, after which the text embeddings are further injected via cross-attention in the spatial transformer. More specifically, the same text embedding is replicated L times and used for all frames.
[0009] The object of the present invention is to provide an optimized and / or more precise method for text-to-video generation, in particular for the generation and enhancement of training and / or test data.
[0010] This object is achieved by a method according to the features of claim 1. This object is also achieved by a device according to the features of claim 10. Summary of the Invention
[0011] According to a first aspect, a method for generating synthetic video data from text prompts is provided, in particular for providing video data for training and / or testing and / or verifying and / or validating a machine learning model, the method comprising:
[0012] - providing an input text prompt describing the content of the video data to be generated;
[0013] - Decompose the provided text prompt into at least two text sub-prompts through a large language model;
[0014] - generating a text embedding for each of at least two text sub-cues; and
[0015] -Based on the generated text embeddings, synthetic video data is generated through a video diffusion model.
[0016] It should be understood that the steps according to the present invention and other optional steps do not necessarily have to be implemented according to the sequence shown, and may also be implemented according to different sequences. Further, other intermediate steps may be provided. In addition, a single step may include one or more sub-steps without therefore departing from the scope of the method according to the present invention.
[0017] According to a second aspect, a device for generating synthetic video data from textual prompts is provided, in particular for providing video data for training and / or testing and / or verifying and / or validating a machine learning model, the device comprising an evaluation and computation unit configured to perform the following steps:
[0018] - providing an input text prompt describing the content of the video data to be generated;
[0019] - Decompose the provided text prompt into at least two text sub-prompts through a large language model;
[0020] - generating a text embedding for each of at least two text sub-cues; and
[0021] -Based on the generated text embeddings, synthetic video data is generated through a video diffusion model.
[0022] Statements made with respect to the program apply correspondingly to the apparatus. It will be understood that language modifications of features formulated with respect to the program may be reformulated with respect to the apparatus in accordance with standard language practice without the need to explicitly set forth such formulations here.
[0023] Despite the architectural changes compared to the T2I model, the current T2V model still uses the same single text prompt for all frames. This is suboptimal for video generation because visual processes can have different visual states, which should be described separately. For example, returning to the text prompt "A boy is getting older" mentioned in the background art, this aging process includes a series of visual changes, such as from "a little boy" to "a middle-aged man" and finally to "an old man". In order to better synthesize this visual process, the proposed method uses a large language model (LLM), such as ChatGPT, to transform the original / provided single text prompt into several descriptive text sub-cues, preferably at least two descriptive text sub-cues. Each text sub-cue preferably describes a specific visual state of the video to be generated. In other words, the method uses the general knowledge of the large language model (LLM) to recapitulate the simple input text prompt by decomposing it into highly descriptive sub-cues. With the help of the LLM, a more fine-grained description of sequential visual states in the video data can be obtained. More specifically, the contextual learning capability of LLM is utilized, a specific example is used to guide LLM, and the text prompt style is specifically explained.
[0024] An example of guidance is as follows: a text prompt "A boy is getting old" is provided for synthetic video data generation. The LLM can then be asked or internally process the following question: "Can you separate the aging process and describe the states of aging separately?" A further restriction of the LLM is that each state will only be described in one sentence, and the coherence between the sub-cues must be considered. In addition, the LLM should generate sub-cues directly without using a narrative style. The output sub-cues may look like, for example, "a little boy" to "a middle-aged man" and "an old man". As another example, for the text prompt "a woman is getting fat", the LLM can divide this initial input into two states, such as "a thin woman" and "a fat woman". Of course, the number of states is not limited to two.
[0025] After obtaining the sub-cues from LLM, the method further derives / generates text embeddings corresponding to each sub-cue.
[0026] A large language model (LLM) is an artificial intelligence (AI) program designed to understand, generate, and use human language. It's "large" because it's built on vast amounts of text data and involves millions or even billions of parameters, which it uses to make predictions about language. These models are trained on a diverse range of internet text sources. The training process involves showing the model examples of text and teaching it to predict the next word in a sentence based on previously presented words. Over time, through a process known as machine learning, the model becomes better at making these predictions. This ability to predict the next word enables LLMs to generate coherent and contextually relevant text, translate language, summarize text, answer questions, and perform many other language-related tasks. One of the key features of large language models is their ability to learn "few shots" or "zero shots," in which they can perform tasks for which they were not explicitly trained simply by understanding instructions given in natural language. This flexibility makes them extremely powerful tools for a wide range of applications, from writing assistance to customer service automation. The GPT (Generative Pretrained Transformer) model developed by OpenAI is a prominent example of a large language model. They have attracted significant attention for their ability to generate human-like text and perform a wide variety of language understanding and generation tasks with high proficiency.
[0027] Video diffusion models (VDMs) are an extension of diffusion models applied to video data generation and processing. Originally developed for tasks like image generation and editing, diffusion models are a class of generative models that learn to create data similar to the distribution of a training set. These models work by gradually transforming random noise samples into structured outputs (such as images or videos) through a series of steps or iterations, effectively "denoising" the input. In the context of video, video diffusion models learn to generate or manipulate video sequences rather than static images. This involves understanding and modeling the complexity of temporal dynamics and visual consistency across frames, which is significantly more challenging than image generation due to the additional dimension of time. VDMs must not only learn the appearance of objects in each frame, but also how they move and change over time, maintaining coherence and continuity across the sequence of frames that make up a video. Video diffusion models are trained on large datasets of video clips and learn to predict the next frame in a sequence given the previous frame, or to generate an entire sequence from a noise distribution. These models can be used for a variety of applications, including: video synthesis, where new video clips are generated from scratch; video editing, such as changing the content of existing videos; and video super-resolution, where the quality of video clips is enhanced.
[0028] Generating a text embedding for each of at least two text sub-cues is the process of transforming a specific piece of text (called a sub-cue) into a numerical representation (called an embedding). This process is the foundation of natural language processing (NLP) and machine learning, enabling computers to understand and use human language.
[0029] A "text embedding" is a high-dimensional vector that captures the semantic meaning, syntactic structure, and context of a text. When generating embeddings for text sub-cues, these discrete pieces of text are effectively mapped into a continuous vector space. Each dimension of the embedding vector represents a latent feature of the text, typically learned from a large dataset during language model training. These are smaller, distinct parts of the larger text prompt or query. For example, if the main prompt is about describing a scene in a forest, the sub-cues can focus on specific aspects like the type of tree, the weather, or the time of day. This involves using a pre-trained language model to process the sub-cues. The LLM analyzes the text and outputs a fixed-size vector for each sub-cue. These vectors serve as a numerical representation of the sub-cue's content and context.
[0030] The present invention focuses on processing and / or enhancing sensor data and focuses on classification, object detection and / or semantic segmentation model training tasks. The video data generated by this method can be used to enhance the training of machine learning models to identify various elements in the data, such as traffic signs, road surfaces, pedestrians, vehicles and other object classes (like trees and sky). Additionally, it can enhance the performance of machine learning models to perform video and audio analysis for regression tasks, enabling them to determine continuous values such as distance, speed, acceleration, and track objects in video data. This is particularly accomplished by providing a strategy for generating more accurate training data for such machine learning models. The present invention can be used to enhance and / or generate video data that is similar to video data derived from radar sensors, lidar sensors, ultrasonic sensors, motion sensors and / or thermal sensors.
[0031] The present invention can be used as a method for efficiently selecting and transferring (training) data from a technical system to a back-end computer, with the goal of reducing data traffic. This selective data transfer is primarily intended to enhance machine learning (ML) systems by providing data for training, testing, verifying, and validating these ML systems. The present invention can also be used as an upstream component in a machine learning tool chain, emphasizing its role in preparing ML systems. The present invention can be used to generate enhanced training data and / or enhanced testing and / or verification and / or validation data. Once the ML system is trained using the enhanced data generated by the present method and apparatus, the ML system can be applied to downstream tasks that demonstrate enhanced performance.
[0032] In another aspect, the text sub-cues decompose the text cues to describe sequential visual states of content to be generated as video data, wherein the sequential visual states of the content are to be represented by at least two frames in the generated video data.
[0033] "Text sub-cues" refers to the breaking down of larger, more complex text cues into smaller, more manageable segments. These sub-cues are designed to capture specific details or aspects of a scene or content for visualization in a video. Sub-cues are used to describe different moments or states in a sequence that unfolds over time. Each state represents a specific point in the narrative or visual progression of the content to be depicted in the video. This sequential nature is crucial for generating content with a coherent flow, similar to how a story progresses or how a scene changes over time. The purpose of breaking a text cue into sub-cues is to create video data. This means transforming the text description into a visual representation that constitutes a video frame. The reference to "at least two frames" emphasizes that the generated video data will be composed of multiple frames, which are necessary to depict the sequential visual states described by the sub-cues. A single frame can represent a static image, but multiple frames are required to convey motion, change, or the progression of time, which are essential elements of video content.
[0034] In another aspect, the method further includes interpolating between two adjacent generated text embeddings to derive a text embedding for another intermediate frame of the video data to be generated.
[0035] "Two adjacent generated text embeddings" preferably refer to embeddings corresponding to two uninterrupted frames or scenes in a video. These frames are "adjacent" in the sense that they are immediately adjacent to each other in the video sequence, representing uninterrupted moments or states in a narrative or visual progression. Interpolation preferably involves calculating or generating an intermediate value between two adjacent text embeddings. This mathematical process is preferably intended to create a smooth transition between the semantic content represented by the two embeddings. By interpolating between these embeddings, the method generates new embeddings that represent intermediate states or scenes that logically fit between the two original frames. The result of the interpolation process is a new text embedding that represents the content of the intermediate frame in the video. This frame is not directly described by the original text cue, but is generated to create a smoother visual or narrative transition in the video content. Thus, interpolation is performed between the generated text / language embeddings, preferably between each state. The interpolated text embedding is preferably fed as a conditional text embedding for each frame. In this way, the text-to-video model has a more explicit guidance about the visual states in the temporal dimension and can therefore better generate the desired video.
[0036] In another aspect, the video diffusion model comprises a convolutional layer, at least one spatial transformer, and at least one temporal transformer, wherein the method further comprises:
[0037] - Extract attention maps from temporal transformers;
[0038] -provide incremental attention maps for regularizing the extracted attention maps; and
[0039] -Regularize the attention map based on the incremental attention map.
[0040] In the context of video diffusion models, an "attention map" refers to a mechanism or visualization that shows how different parts of a video frame or sequence are weighted or focused on by the model during processing or generation tasks. Video diffusion models are an extension of diffusion models applicable to video generation, relying on capturing both spatial and temporal dependencies in video data to efficiently generate or manipulate video sequences. The attention mechanism in these models allows the network to focus on specific parts of the input data (in this case, video frames or text embeddings) more than others when performing the task. This enables the VDM to prioritize relevant spatial features (e.g., objects, shapes) and temporal features (e.g., motion, changes over time) that are critical for understanding scenes and generating coherent video content. An attention map is essentially a visualization or representation of these focus areas, showing where the model is "focusing" at any given time during video processing. It can highlight spatial attention—that is, naming which regions within a single frame are focused on. This can help understand how the model perceives different objects and elements in the scene. Furthermore, it can highlight temporal attention—that is, how the model's focus shifts across frames. This can reveal how the model tracks motion and changes over time, which is crucial for capturing the dynamic nature of video content.
[0041] The provided method also shows a measure to regularize the temporal attention layer to further enforce the desired visual patterns. Although the former VDM focuses too much on specific frames, such as the initial frame, and thus typically generates video data consisting of only similar frames, the provided aspect alleviates this problem by defining an additional incremental attention map, which is added to and / or processed together with the (former) temporal attention map. By adopting this regularization, some prior information of the desired visual patterns is included, and thus more attractive results are achieved in the generated video data. Therefore, in addition to re-interpretation, a temporal attention regularization strategy is also provided. The temporal attention map can be visualized because temporal attention plays a vital role in modeling the temporal relationship between frames of video data. As input to the VDM temporal transformer, the features are first rearranged into (h×w, L, c), where the spatial dimension is regarded as the batch dimension. Therefore, the temporal attention is preferably self-attention between L frames. These intermediate features are then projected into query (Q), key (K) and value (V). The attention map is then preferably obtained by calculate.
[0042] After reshaping, the attention map A preferably has R L×LThe shape of , where L is the number of frames. The attention map created / generated in this way may not be very structured because usually too much attention is given to the initial frame rather than to the adjacent frames. Therefore, despite the application of various cue conditions, the VDM tends to synthesize frames similar to the initial state, causing convergence of similar appearance. This shortcoming may come from the training process of the VDM, where the training video clips are usually short and have limited dynamic movement, so the trained VDM may pay too much attention to specific frames. To alleviate this shortcoming, the present method can regularize the temporal attention map when inferring time. More specifically, in order to encourage the model to synthesize more drastic visual changes in the video data, an incremental attention map ΔA can be defined. The incremental attention map can be defined manually. Therefore, the attention distribution along the diagonal offset direction can be modeled as a Gaussian distribution, that is
[0043]
[0044] where k∈{0,1,…,L2} is the offset distance from the diagonal and σ is the standard deviation.
[0045] On the other hand, regularizing the attention map involves:
[0046] - transform the extracted attention map and the incremental attention map based on an affine transformation; or
[0047] -Add the extracted attention map and the incremental attention map multiplied by the scaling factor.
[0048] The updated attention map can be formulated as a transformation based on the original attention map A and the delta map ΔA:
[0049] A′←F(A, ΔA),
[0050] Where F represents the affine transformation.
[0051] Alternatively, a summation operation can be performed, i.e.
[0052] A′←A+αΔA,
[0053] where α is the scaling factor.
[0054] The incremental attention map can also be created automatically, i.e. based on a pre-trained machine learning model. The incremental attention map can have a diagonal matrix-like representation. However, other alternative incremental attention maps can be used to implement different video modes, such as looping videos.
[0055] In another aspect, providing an incremental attention map to regularize the extracted attention map includes:
[0056] - We obtain incremental attention maps from a reference video by computing the correlation between its visual features extracted from a pre-trained visual encoder (specifically the CLIP image encoder).
[0057] If a reference video is given, the incremental attention map can be obtained by computing the correlation between visual features extracted from a pre-trained visual encoder (e.g., CLIP image encoder). It is noted that different incremental attention matrices can be applied at different resolutions to achieve finer-grained control. Attention at different resolutions can control appearance and / or color and / or style and / or content and / or layout respectively. Therefore, the attention map can be manipulated separately at different resolutions for style and / or content changes in the video data, depending on what the user wants to generate.
[0058] In another aspect, providing an incremental attention map to regularize the extracted attention map includes:
[0059] - An adaptive network is trained to extract learnable incremental attention maps and affine transformations to capture dynamic patterns in video data and learn attention regularization.
[0060] In addition to direct regularization, when paired text-video data are available, one can train an adaptive network f to learn attention regularization effects, such as learnable incremental attention matrices and affine transformations to better capture dynamic patterns:
[0061] A′=f(Wp,V,A),
[0062] where Wp and V represent text embedding and training video respectively.
[0063] In summary, through the provided approach, two main strategies, namely, re-narration and temporal attention regularization, are proposed to improve the pre-trained text-to-video diffusion model, leading to attractive visual results that well conform to the input prompts.
[0064] In another aspect, a computer program is provided, comprising program code, which, when executed on a computer, performs at least part of the method according to one of the embodiments of the present invention. In other words, according to the present invention, a computer program (product) comprises instructions, which, when executed by a computer, cause the computer to perform the method / steps according to one of the embodiments of the present invention.
[0065] In another aspect, a computer-readable medium is provided, comprising the program code of a computer program, which is also provided, for carrying out, when the computer program is executed on a computer, at least part of the method according to the present invention in one of its embodiments. In other words, the present invention relates to a computer-readable (storage) medium comprising instructions which, when executed by a computer, cause the computer to carry out the method / steps of the method according to one of its embodiments.
[0066] The described embodiments and further developments can be combined with one another as desired.
[0067] Further possible embodiments, further developments and implementations of the invention also include combinations of features of the invention described above or below with reference to exemplary embodiments not explicitly mentioned. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings are intended to provide a further understanding of the embodiments of the present invention. They illustrate the embodiments and, together with the description, serve to explain the principles and concepts of the present invention.
[0069] Other embodiments and many of the advantages mentioned will be apparent with respect to the accompanying drawings.The elements shown in the drawings are not necessarily shown to scale with respect to each other.
[0070] In the accompanying drawings:
[0071] Figure 1 is a schematic flow chart of a method for generating synthetic video data based on text prompts;
[0072] FIG2 is a video diffusion model according to the prior art;
[0073] Figure 3 is a block diagram of a method for generating synthetic video data based on textual prompts;
[0074] Figure 4 is a schematic representation of the regularized attention map;
[0075] Figure 5 is the alternative incremental attention map; and
[0076] Figure 6 is an example of generated video content consisting of 16 frames.
[0077] Throughout the figures of the drawings, unless otherwise indicated, like reference numerals designate identical or functionally identical elements, parts or components. DETAILED DESCRIPTION
[0078] Figure 1A schematic flow chart of a method for generating synthetic video data based on text prompts is shown, in particular for providing video data for training and / or testing and / or verifying and / or validating a machine learning model.
[0079] In any embodiment, the method may be at least partially performed by the device 100. For this purpose, the device 100 may include several components not shown in greater detail, such as one or more supply units and / or at least one evaluation and calculation unit. It should be understood that the supply device may be formed together with the evaluation and calculation unit or may be formed separately therefrom. Furthermore, the device may include a storage unit and / or an output unit and / or a display unit and / or an input unit (in particular for providing text prompts).
[0080] According to the present invention, the computer-implemented method comprises at least the following steps:
[0081] In step S1 , the method comprises providing an input text prompt describing the content of video data to be generated.
[0082] In step S2 , the method includes decomposing a provided text prompt into at least two text sub-prompts by means of a large language model.
[0083] In step S3, the method includes generating a text embedding for each of the at least two text sub-cues.
[0084] In step S4, the method includes generating synthetic video data through a video diffusion model based on the generated text embeddings.
[0085] FIG2 shows a video diffusion model (VDM) 200 according to the prior art. The video diffusion model 200 includes a convolutional layer 202, a spatial transformer 204, and a temporal transformer 206. The video data generated by the VDM 200 is displayed on the time axis z. t to z t-1 For each frame along the time axis, a representation of the VDM 200 is presented. In the prior art, each frame processed by the VDM 200 results in a text hint y 208, which can be transformed into a text embedding 210 as input. The transformation to text embedding can be done by CLIP 文本 The model is complete. In the prior art, the same text prompt is provided for each frame.
[0086] Figure 3 The method according to the present invention is shown. In which the text prompt 300 initially provided is decomposed into at least two (in Figure 3For each of at least two text sub-prompts 302, a text embedding 306 is generated. The transformation / generation of the corresponding text embedding 306 can be performed by CLIP 文本 Based on the generated text embedding 306, synthetic video data 308 is generated by the pre-trained video diffusion model 310. Figure 3 , interpolation is performed between two adjacent generated text embeddings 306 to derive an interpolated text embedding 312 of another intermediate frame of the video data to be generated. The interpolation may be performed on L frames.
[0087] Figure 4 Regularization of an attention map 400 (also referred to as an attention matrix) is shown. As shown in FIG2 , the video diffusion model 200 , 310 comprises a convolutional layer, at least one spatial transformer and at least one temporal transformer. Attention map(s) 400 are extracted from the temporal transformer, preferably for each frame or between every two adjacent frames. Furthermore, an incremental attention map 402 is provided for regularizing the extracted attention map 400. The incremental attention map 402 may be provided manually. The attention map(s) 400 (each) is regularized based on the incremental attention map 402 to obtain a regularized attention map(s) 404. In Figure 4 , the incremental attention map 402 has the shape of a diagonal matrix.
[0088] Figure 5 Additional possible incremental attention maps 402 are shown. One alternative incremental attention map 500 shows a V-shape and represents regularization of video content that shows a cycle, i.e., a young person gets older and then gets younger again. Another alternative incremental attention map 502 shows a parallel stripe shape and represents regularization of video content that shows a repeating pattern, i.e., a young person gets older and then a young person gets older.
[0089] like Figure 5 As also shown, by calculating 601 the correlation between visual features 603 of each frame of a reference video 602 extracted from a pre-trained visual encoder 604 (particularly a CLIP image encoder), an incremental attention map 600 is obtained from the reference video 602, which can provide an incremental attention map 402 for regularization of the extracted attention map 400.
[0090] Figure 6 An example of generated video data content consisting of 16 frames is shown. The generated synthetic video data shows a boy aging. The proposed method, which generates videos based on sub-cues and regularizes the attention map with an incremental attention map, enhances the accuracy of the generated content.
Claims
1. A method for generating synthetic video data from textual prompts, in particular for providing video data for training and / or testing and / or verifying and / or validating a machine learning model, the method comprising: - providing (S1) an input text prompt describing the content of the video data to be generated; - decomposing (S2) the provided text prompt into at least two text sub-prompts by a large language model; - generating (S3) a text embedding for each of the at least two text sub-cues; and -Based on the generated text embeddings, synthetic video data is generated (S4) through a video diffusion model.
2. The method of claim 1 , wherein the text sub-cues decompose the text cues to describe sequential visual states of the content to be generated as the video data, wherein The sequential visual states of the content will be represented by at least two frames in the generated video data.
3. The method of claim 1 or claim 2, wherein the method further comprises interpolating between two adjacent generated text embeddings to derive a text embedding for another intermediate frame of the video data to be generated.
4. A method according to any one of the preceding claims, wherein The video diffusion model comprises a convolutional layer, at least one spatial transformer and at least one temporal transformer, wherein the method further comprises: - extracting an attention map from the temporal transformer; -provide incremental attention maps for regularizing the extracted attention maps; and -regularizing the attention map based on the incremental attention map.
5. The method of claim 4, wherein regularizing the attention map comprises: - transforming the extracted attention map and the incremental attention map based on an affine transformation; or -Add the extracted attention map and the incremental attention map multiplied by a scaling factor.
6. The method of claim 4, wherein providing the incremental attention map to regularize the extracted attention map comprises: - Obtaining the incremental attention map from a reference video by computing correlations between visual features of the reference video extracted from a pre-trained visual encoder, in particular the CLIP image encoder).
7. The method of claim 4, wherein providing the incremental attention map to regularize the extracted attention map comprises: - An adaptive network is trained to extract learnable incremental attention maps and affine transformations to capture dynamic patterns in the video data and learn attention regularization.
8. A computer program comprising program code for carrying out at least part of the method according to any one of claims 1 to 7 when the computer program is executed on a computer.
9. A computer-readable medium containing program code of a computer program for executing at least part of the method according to any one of claims 1 to 7 when the computer program is executed on a computer.
10. A device (100) for generating synthetic video data from textual prompts, in particular for providing video data for training and / or testing and / or verifying and / or validating a machine learning model, the device (100) comprising an evaluation and computation unit configured to perform the following steps: - providing an input text prompt describing the content of the video data to be generated; - Decompose the provided text prompt into at least two text sub-prompts through a large language model; - generating a text embedding for each of the at least two text sub-cues; and -Based on the generated text embeddings, synthetic video data is generated through a video diffusion model.