Windowed attention for text-to-video generation

CN122534293APending Publication Date: 2026-08-07ADOBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ADOBE INC
Filing Date
2026-02-06
Publication Date
2026-08-07

Smart Images

  • Figure CN122534293A_ABST
    Figure CN122534293A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to windowed attention for text-to-video generation. A method, apparatus, non-transitory computer-readable medium, and system for video generation includes obtaining an input prompt that describes a scene. A video generation model generates global labels based on the input prompt by performing an attention process based on a plurality of frame labels. The video generation model generates frame labels by performing an attention process based on the global labels and a subset of the plurality of frame labels in a local window. Subsequently, the video generation model generates a synthetic video based on the frame labels, where the synthetic video depicts the scene and includes image frames corresponding to the frame labels.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application is based on U.S. Patent Application No. 63 / 755461, filed February 7, 2025, with the United States Patent and Trademark Office, and claims priority under 35 USC §120, the contents of which are incorporated herein by reference in their entirety. Background Technology

[0002] The following generally deals with machine learning, and more specifically, video generation using machine learning models. Machine learning algorithms build models based on sample data (called training data) to make predictions or decisions in response to inputs, without needing to be explicitly programmed to do so. One application area of ​​machine learning is video generation.

[0003] For example, machine learning models can be trained to predict video features in response to input cues, and then generate videos based on those predicted features. In some cases, cues can be used to perform complex video manipulations and compositions. This video generation enables users to edit videos and generate modified videos with desired features, thus making video generation easier for laypeople. Summary of the Invention

[0004] This disclosure describes systems and methods for video generation. Embodiments of this disclosure are configured to generate high-resolution videos based on input prompts. In some cases, the video generation model of this disclosure is configured to implement a windowed self-attention mechanism in a video-based diffusion network. According to one embodiment, the video generation model is configured to perform a global update followed by a local update. For example, the video generation model updates register tags, video tags, and text tags based on input prompts and a set of video frames associated with the video.

[0005] A method, apparatus, and non-transitory computer-readable medium for video processing are described. One or more aspects of the method, apparatus, and non-transitory computer-readable medium include: obtaining an input cue describing a scene; generating a global marker based on the input cue by performing an attention process based on multiple frame markers using a video generation model; generating frame markers by performing an attention process based on the global markers and a subset of multiple frame markers in a local window using the video generation model; and generating a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0006] A method, apparatus, and non-transitory computer-readable medium for video processing are described. One or more aspects of the method, apparatus, and non-transitory computer-readable medium include: obtaining an input cue describing a scene; generating text markers based on the input cue; generating frame markers using a video generation model by performing an attention process based on the text markers and a subset of frame markers in a local window; and generating a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0007] An apparatus and system for video processing are described. One or more aspects of the apparatus and system include: obtaining an input cue describing a scene; generating a global marker based on the input cue by performing an attention process based on multiple frame markers using a video generation model; generating frame markers by performing an attention process based on the global markers and a subset of multiple frame markers in a local window using the video generation model; and generating a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers. Attached Figure Description

[0008] Figure 1 An example of a video processing system according to various aspects of this disclosure is shown.

[0009] Figure 2 Examples of methods for video generation according to various aspects of this disclosure are shown.

[0010] Figure 3 An example of a video generation process according to various aspects of this disclosure is shown.

[0011] Figure 4 An example of a video generation model based on various aspects of this disclosure is shown.

[0012] Figure 5 An example of a converter architecture according to various aspects of this disclosure is shown.

[0013] Figure 6 Examples of potential diffusion processes according to various aspects of this disclosure are shown.

[0014] Figure 7 An example of a diffusion converter network according to various aspects of this disclosure is shown.

[0015] Figure 8 An example of a diffusion converter network for video is shown according to various aspects of this disclosure.

[0016] Figure 9 An example of a denoising diffusion process according to various aspects of this disclosure is shown.

[0017] Figure 10 Examples of methods for video processing according to various aspects of this disclosure are shown.

[0018] Figure 11 Examples of methods for video processing according to various aspects of this disclosure are shown.

[0019] Figure 12 Examples of methods for video processing according to various aspects of this disclosure are shown.

[0020] Figure 13 Examples of methods for video processing according to various aspects of this disclosure are shown.

[0021] Figure 14 Examples of methods for training machine learning models according to various aspects of this disclosure are shown.

[0022] Figure 15 Examples of methods for training diffusion models according to various aspects of this disclosure are shown.

[0023] Figure 16 Examples of computing devices according to various aspects of this disclosure are shown.

[0024] Figure 17 Examples of video processing apparatuses according to various aspects of this disclosure are shown. Detailed Implementation

[0025] Existing systems use diffusion networks for video generation. In some cases, existing video generation models use transformer networks with self-attention mechanisms. For example, such systems can incorporate spatial information into the generated video. However, when the generated video has high resolution, such systems cannot capture both global and spatial relevance in the generated video.

[0026] Furthermore, in some cases, existing video generation systems become scalable as the resolution of the generated video increases. In some situations, existing systems become computationally challenging when processing real-world scenes using long-duration, high-resolution videos. As a result, such systems have limited effectiveness and tend to depict poor temporal coherence and spatial detail in the generated videos.

[0027] In contrast, embodiments of this disclosure are configured to generate high-resolution videos based on input text prompts. In some cases, the video generation model of this disclosure is configured to implement a windowed self-attention mechanism in a video-based diffusion network. According to one embodiment, the video generation model is configured to perform global and local updates. For example, the video generation model updates register tags, image tags, video tags, and text tags based on the input prompts and a set of video frames associated with the video.

[0028] This disclosure describes systems and methods for video generation. Embodiments of this disclosure are configured to perform a windowed attention mechanism for a diffusion network. In some cases, the video generation model of this disclosure is based on a diffusion network. In some cases, the video generation model is configured to perform segmentation of a video sequence. For example, the video generation model performs segmentation to generate one or more non-overlapping temporal videos. In some examples, the size or duration of the non-overlapping temporal videos is smaller than the size or duration of the input video sequence.

[0029] In some cases, video generation models involving diffusion networks are configured to perform a self-attention mechanism on each of one or more non-overlapping temporal videos. In other cases, the video generation model is able to capture fine-grained temporal dynamics of each of one or more non-overlapping temporal videos.

[0030] Embodiments of this disclosure are configured to perform a non-overlapping windowed attention mechanism. In some cases, the video generation model of this disclosure modifies the window size. For example, the video generation model reduces the complexity associated with the window size based on a linear change. In some examples, the video generation model performs segmentation of the input video sequence based on a variable window size. In some examples, the window size at each attention layer of the video generation model is variable.

[0031] In some cases, embodiments of this disclosure are used to optimize the efficiency and performance of video generation models by varying the window size at each attention layer. Additionally, in other cases, variations in the window size at each attention layer are used to achieve efficient processing of high-resolution videos.

[0032] One embodiment of this disclosure is configured to perform video generation. In some cases, a video generation model is used to learn global tags and global context. For example, the video generation model learns global tags from keyframes associated with each of one or more non-overlapping time videos. For example, the video generation model learns global context from keyframes associated with each of one or more non-overlapping time videos. By using a video generation model to learn global tags and global context from keyframes, embodiments of this disclosure can efficiently propagate the global context used for video generation.

[0033] In some cases, global tags are considered as a compact representation of the input video sequence. For example, global tags capture long-range correlations based on the input video sequence. Accordingly, by combining learnable global tags from keyframes associated with each of one or more non-overlapping temporal videos, embodiments of this disclosure are able to capture long-range correlations while reducing the computational cost associated with self-attention mechanisms.

[0034] In some cases, the context of keyframes associated with each of one or more non-overlapping temporal videos ensures temporal consistency. For example, the keyframe context ensures temporal consistency across windows at each attention layer. Accordingly, by capturing the keyframe context to obtain temporal consistency across windows, embodiments of this disclosure are able to prevent artifacts in the generated video. Additionally, by ensuring temporal consistency, embodiments are able to prevent incoherent transitions between sets of video frames.

[0035] Embodiments of this disclosure are configured to perform a video generation process. In some cases, the video generation model of this disclosure is configured to vary the window size at different attention layers of the diffusion network. For example, by varying the window size at different attention layers, embodiments of this disclosure can optimize the efficiency and performance of video generation.

[0036] One embodiment of this disclosure is configured to implement learnable register tagging during the video generation process. In some cases, the video generation model uses learnable register tagging (i.e., global tagging) to propagate global context between windows at each attention layer of the diffusion network. In other cases, the video generation model is configured to generate synthetic videos based on a diffusion network architecture.

[0037] One embodiment of this disclosure is configured to combine context from keyframes associated with each of one or more non-overlapping temporal videos. In some cases, the video generation model is used to implement context from nearby keyframes. For example, nearby keyframes refer to keyframes (or multiple keyframes) adjacent to the current window. In other cases, the video generation model is used to implement context from a moving window that moves across an attention layer. For example, the window moves across an attention layer of a video generation model that includes a diffusion network. By combining context from nearby keyframes and moving windows, embodiments of this disclosure ensure the consistency and efficient propagation of global context across the diffusion network architecture of the video generation model.

[0038] As used in this paper, input cues are text-based or natural language instructions that guide or condition the video generation process. For example, an input cue describes the scene a user wants to depict in a synthetic video. In some cases, input cues are encoded using a text encoder to generate tokens that act as conditional input signals to downstream modules, such as video generation models. The video generation model can use these cues to determine the scene composition, object placement, temporal progression, or stylistic attributes of the generated content.

[0039] In some cases, transformer-based diffusion models operate by performing attention operations on tags, which are sequences of vectors representing the video being generated. In other cases, tags represent patches of an image, where each patch is iteratively denoised to generate image content, and after denoising, the patches are reconstructed to form the final video. In still other cases, some tags may require fewer processing operations than others during the iterative generation of the video.

[0040] As described in this article, a keyframe in a video clip refers to an important frame in a sequence of video frames. Important frames serve as references for frame interpolation while ensuring temporal coherence. For example, a keyframe can refer to the first frame in a frame block, and keyframes include block information, such as, but not limited to, pixel information.

[0041] In some cases, a frame tag is a data representation that identifies or encodes frames within a sequence of video frames. In other cases, a frame tag is used as a reference point for subsequent frames. For example, a frame tag includes motion information associated with video frames within a block.

[0042] In some cases, a global tag refers to a proxy tag. In others, a global tag may not be part of the video but can refer to additional parameters to be optimized in the network. In still others, global tags are used to perform message passing. In some examples, a global tag can refer to a register tag that includes information about different input tags. For instance, a global tag can include a learnable embedding that interacts with the input tags during processing via an attention mechanism. Global tags exert attention on other tags (e.g., text tags, frame tags, etc.) and are also subject to attention from other tags (e.g., text tags, frame tags, etc.), thereby accumulating contextual information across the sequence.

[0043] As used in this paper, a video generation model can refer to a machine learning-based model configured to generate media items (such as videos) based on input data. This model can include machine learning architectures such as neural networks, generative adversarial networks (GANs), transformer-based models, or diffusion networks. In some cases, video generation models can leverage trained parameters, latent representations, or probabilistic sampling techniques to generate outputs that resemble or expand upon the input data while maintaining desired style or structural properties.

[0044] As used herein, an attention process refers to a mechanism for determining how much influence each token in a sequence should have on the representation of another token when computing context embeddings. For example, in some cases, each token is first projected into three vectors: query (Q), key (K), and value (V). The model computes the similarity between the query of one token and the key of another token (e.g., using a scaled dot product), applies softmax to obtain normalized attention weights, and then forms a weighted sum of the value vectors. This process allows the model to dynamically focus on the most relevant parts of the input at each location, thereby capturing relevance across arbitrary distances in the sequence. In some embodiments, multiple attention heads perform this process in parallel with different learned projections, allowing the transformer to capture multiple relational patterns simultaneously. According to embodiments of this disclosure, the scope of a token's influence may be influenced by the tokens used in the attention process. For example, some tokens may learn representations based on global context, while others may learn representations based on more local context (e.g., by using fewer tokens in the attention process).

[0045] As described in this paper, synthetic video refers to a temporally ordered sequence of frames or images. In some cases, synthetic video depicts a temporally coherent sequence of image frames synthesized based on input cues. Synthetic video reflects both the spatial structure and temporal dynamics generated based on video generation models, such as motion continuity, object persistence, and scene transitions.

[0046] Embodiments of this disclosure improve the efficiency of traditional machine learning models. For example, by segmenting a video into a set of non-overlapping videos, embodiments of this disclosure achieve localized computation of attention information, resulting in a significant reduction in computational overhead. Furthermore, by performing a windowed attention mechanism on non-overlapping temporal videos to learn global context, embodiments of this disclosure are able to capture fine-grained spatial information of the generated video, leading to enhanced generation efficiency, scalability, and resolution quality. By combining windowed attention mechanisms, global tagging, and keyframe context, embodiments of this disclosure provide a scalable and efficient framework for video generation while enhancing computational efficiency and video quality.

[0047] Embodiments of this disclosure can be implemented in a video processing system. For example, a video processing system based on this disclosure acquires input prompts (e.g., input prompts describing a target action) and generates an output video that accurately depicts the target action described in the input prompt. Reference Figures 1 to 3 A sample application for generating videos depicting elements is provided. (Reference) Figures 4 to 9 and Figures 16 to 17 Details about the architecture of the video generation model are provided. (Reference) Figures 10 to 13 Details regarding the operation of the video generation model are provided. (Reference) Figures 14 to 15An example of the process for training a video generation model is provided. Video generation system

[0048] refer to Figures 1 to 9 Systems and apparatuses for video processing are described. Figure 1 An example of a video processing system 100 according to various aspects of this disclosure is shown. In one aspect, the video processing system 100 includes a user 105, a user device 110, a video processing apparatus 115, a cloud 120, and a database 125.

[0049] exist Figure 1 In one example, the user provides a prompt with an action (e.g., running) to the video processing device 115 via a user interface provided by the video processing device 115 on the user device 110. In some examples, the input prompt is input text (such as...). Figures 1 to 2 As shown). Figure 1 As shown, the input prompt is text that provides details about elements and actions (e.g., "a dog running in the street"), where the user wants to use the video processing apparatus 115 of this disclosure to generate a video based on that text.

[0050] In some cases, the video processing device 115 implements a video generation model (such as at least referencing...). Figures 4 to 9 The described video generation model generates videos based on input prompts. In some cases, such as... Figure 1 As shown, a user provides an input prompt (e.g., a text query) to the video processing device 115, where the user wants to depict aspects of the input prompt in the video. In some examples, the video processing device generates a video that is precisely aligned with the information provided by the input prompt.

[0051] In some examples, the video processing apparatus generates spatially consistent videos that accurately capture fine-grained temporal dynamics based on efficient propagation of the global context during video generation. Video processing apparatus 115 is a reference... Figure 3 Examples of the corresponding elements described, or including aspects thereof.

[0052] Refer again Figure 1For example, video processing device 115 generates video that accurately depicts (or modifies) aspects (e.g., elements) described by input prompts. According to some aspects, user device 110 is a personal computer, laptop computer, mainframe computer, handheld computer, personal assistant, mobile device, or any other suitable processing device. In some examples, user device 110 includes software that displays a user interface (e.g., a graphical user interface) provided by video processing device 115. In some aspects, the user interface enables the transfer of information (such as video (input or generated video), images, prompts, canvases, etc.) between user 105 and video processing device 115.

[0053] According to some aspects, the user equipment user interface enables user 105 to interact with user equipment 110. In some embodiments, the user equipment user interface may include an audio device (such as an external speaker system), an external display device (such as a display screen), or an input device (e.g., a remote control device that interfaces directly with the user interface or via an I / O controller module). In some cases, the user equipment user interface may be a graphical user interface.

[0054] According to some aspects, the video processing apparatus 115 includes a computer-implemented network. In some embodiments, the computer-implemented network includes a video generation model (such as at least referenced...). Figures 4 to 9 The video generation model described herein. In some embodiments, the video processing apparatus 115 further includes one or more processors, a memory subsystem, a communication interface, an I / O interface, one or more user interface components, and a bus, as described in reference. Figure 16 As described. Additionally, in some embodiments, the video processing device 115 communicates with the user equipment 110 and the database 125 via the cloud 120.

[0055] In some cases, the video processing apparatus 115 is implemented on a server. The server provides one or more functions to users linked via one or more various networks, such as a cloud 120. In some cases, the server includes a single microprocessor board containing a microprocessor responsible for controlling all aspects of the server. In some cases, the server uses the microprocessor and protocols to exchange data with other devices or users on one or more networks via Hypertext Transfer Protocol (HTTP) and Simple Mail Transfer Protocol (SMTP), although other protocols such as File Transfer Protocol (FTP) and Simple Network Management Protocol (SNMP) may also be used. In some cases, the server is configured to send and receive files in Hypertext Markup Language (HTML) format (e.g., for displaying web pages). In various embodiments, the server includes a general-purpose computing device, a personal computer, a laptop computer, a mainframe computer, a supercomputer, or any other suitable processing device.

[0056] Cloud 120 is a computer network configured to provide on-demand availability of computer system resources, such as data storage and computing power. In some examples, Cloud 120 provides resources without active user management. The term "cloud" is sometimes used to describe data centers available to many users via the Internet. The functionality of some large cloud networks is distributed across multiple locations originating from a central server. Servers are designated as edge servers if they have direct or close connections to users. In some cases, Cloud 120 is limited to a single organization. In other examples, Cloud 120 is available to many organizations. In one example, Cloud 120 includes a multi-layered communication network comprising multiple edge routers and a core router. In another example, Cloud 120 is based on a cluster of local switches in a single physical location. According to some aspects, Cloud 120 provides communication between user equipment 110, video processing device 115, and database 125.

[0057] Database 125 is a set of organized data. In one example, database 125 stores data in a specified format called a schema. Depending on some aspects, database 125 may be constructed as a single database, a distributed database, multiple distributed databases, or an emergency backup database. In some cases, a database controller manages the data storage and processing within database 125. In some cases, a user interacts with the database controller. In other cases, the database controller operates automatically without user interaction. Depending on some aspects, database 125 is external to video processing device 115 and communicates with video processing device 115 via cloud 120. Depending on some aspects, database 125 is included within video processing device 115.

[0058] Figure 2 Examples of a method 200 for video generation according to aspects of this disclosure are shown. In some examples, these operations are performed by a system including a processor that executes a set of code to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed based on the methods and processes described according to aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.

[0059] According to one embodiment of this disclosure, a video processing apparatus (such as reference ) Figure 3 and Figure 17 The described video processing apparatus provides a video generation model (such as a reference). Figures 4 to 9 and Figure 17 The described video generation model accurately generates synthetic videos that depict the elements described in the input query or the actions performed by the elements described in the input query.

[0060] In step 205, the system provides a query that includes text. In some cases, this step involves referencing... Figure 1 The user described, or the one who can perform the action.

[0061] In some cases, a text query describes the object upon which the user wants to generate a video. Additionally or alternatively, a text prompt provides the action the user wants to generate the composite video based on. For example, the user provides a text prompt instructing a video processing device to generate a composite video precisely aligned with the text prompt. Figure 2 As shown, the user accesses the user device's user interface (such as reference). Figure 1 The user interface of the user equipment 110 described herein provides text prompts to the video processing device, such as “dog running in the street”.

[0062] In operation 210, the system generates video based on the query. In some cases, this step involves referencing... Figure 1 and Figure 3 The video processing apparatus described, or that can be performed by it.

[0063] In some cases, the video processing apparatus includes a video generation model that includes a diffusion model (such as a reference model) for generating synthetic video. Figures 6 to 9 The diffusion model described. In some cases, the video generation model generates synthetic videos that include objects described in text prompts. In other cases, the video generation model generates synthetic videos that include actions described in text prompts. For example... Figure 2 As shown, the video processing device generates a composite video depicting a "dog running in the street" provided by the user in operation 205.

[0064] In some cases, the video generation model of this disclosure is based on a diffusion network. In some cases, the video generation model is configured to perform segmentation of the input video sequence (e.g., input noise) to generate one or more non-overlapping temporal videos. In some cases, the video generation model including the diffusion network performs a self-attention windowing mechanism on each of the one or more non-overlapping temporal videos. In some cases, the video generation model is able to use the self-attention windowing mechanism to capture fine-grained temporal dynamics of each of the one or more non-overlapping temporal videos.

[0065] In step 215, the system provides the generated video to the user. In some cases, this step involves referencing... Figure 1 and Figure 3 The described video processing apparatus, or the apparatus that can perform the processing, generates video via user equipment (such as...). Figure 1 The user interface of the user device 110 described herein is provided to the user. (At least refer to...) Figures 4 to 10 Further details regarding the generation of the video are provided.

[0066] Figure 3 An example of a video generation process 300 according to various aspects of this disclosure is shown. In one aspect, the video generation process 300 includes an input prompt 305, a video processing device 310, and a synthesized video 315.

[0067] In some examples, the input prompt includes a 305 description element. In some examples, such as... Figure 3 As shown, input prompt 305 describes the action performed by the element. In some examples, the user provides input prompt 305 to the video processing device 310 via the user interface of the video processing device 310. Input prompt 305 is a reference Figures 1 to 2 Examples of the corresponding elements described, or including aspects thereof.

[0068] The video processing apparatus 310 disclosed herein (such as at least referenced) Figures 1 to 2 and Figure 17 The described video processing apparatus 310 receives input prompts 305 from a user. In some cases, the video processing apparatus 310 includes a video generation model that includes a video-based diffusion network (such as a reference network). Figures 6 to 9 The described diffusion network) and transformer models for generating synthetic videos based on windowed self-attention methods (such as reference) Figure 5 The described converter model.

[0069] In some cases, windowed self-attention mechanisms are used in video-based diffusion networks. According to one embodiment, the video generation model is configured to perform global and local updates. For example, the video generation model updates register tags, image tags, video tags, and text tags based on input prompts and a set of video frames from the video.

[0070] In some cases, video generation models are configured to perform segmentation of the input video sequence to generate one or more non-overlapping temporal videos. In other cases, video generation models involving diffusion networks are configured to perform a windowed self-attention mechanism on each of the one or more non-overlapping temporal videos.

[0071] In some cases, the video generation model is used to learn global tags and global context. For example, the video generation model learns global tags from keyframes associated with each of one or more non-overlapping temporal videos. By learning global tags and global context from keyframes, the video processing apparatus 310 is able to efficiently propagate the global context used for video generation. The video processing apparatus 310 is a reference... Figure 1 Examples of the corresponding elements described, or including aspects thereof. Video 315 is for reference. Figures 1 to 2 and Figure 4 Examples of the corresponding elements described, or including aspects thereof.

[0072] Figure 4 An example of a video generation process 400 according to various aspects of this disclosure is shown. In one aspect, the video generation process 400 includes a global marker 405, multiple markers 410, a text marker 415, a keyframe 420, a frame marker 425, a local window 430, and a subsequent window 435.

[0073] Embodiments of this disclosure include a video generation model configured to perform windowed self-attention operations. In some cases, the attention operation includes a process of simultaneously updating global and frame tags. In some cases, the attention operation includes a process of updating the global tag (e.g., generating an updated global tag 405-b). Additionally, in some cases, the attention operation includes a process of generating an updated frame tag 425-b and an updated keyframe 420-b.

[0074] For example, global tag 405 (i.e., each of global tag 405-a and the updated global tag 405-b) consists of learnable register tags, image tags, video tags, text tags, or combinations thereof. In some cases, such as references Figure 4 As shown in step 1, the updated global marker 405-b is generated by performing an attention process based on text marker 415-a and each of the plurality of markers 410. For example, the plurality of markers 410 includes a plurality of keyframes (such as keyframe 420) and a corresponding or subsequent set of frame markers (such as frame marker 425). In some examples, the global marker 405-b is generated by updating the global marker 405-a by performing a windowed self-attention operation (as indicated by the arrows applying attention to each marker) on the global marker 405-a, keyframe 420, frame marker 425, and text marker 415-a.

[0075] Furthermore, text marker 415-a is updated by performing a windowed self-attention operation based on global marker 405-a and text marker 415-a to generate an updated text marker 415-b. In some examples, text marker 415-a is based on input prompts (such as references). Figures 1 to 3 It is generated based on the input prompt described.

[0076] In some cases, video generation models utilize local windows to update frame labels. For example, a video generation model performs updates by executing an attention process based on the labels within the current window. Figure 4 As shown in step 2, the video generation model performs local updates by performing an attention process based on each frame tag within the current window (such as local window 430). In some cases, each frame tag within the current window applies attention to the global tag. In other cases, each frame tag includes unique information associated with a frame block.

[0077] In some cases, keyframe 420-a is updated by performing an attention process based on global marker 405, text marker 415, and frame marker 425-a within the local window 430 to generate an updated keyframe 420-b. In other cases, the frame marker is updated by performing an attention process based on keyframe 420-a within the current window and subsequent keyframes. For example, subsequent keyframes refer to keyframes that are part of a subsequent window.

[0078] like Figure 4 As shown, frame markers (such as each frame marker 425-a of local window 430) apply attention to subsequent keyframes of subsequent local windows. Figure 4 Step 2 in the diagram shows arrows to depict attention to each keyframe marker (such as keyframe 420) corresponding to local window 430. Accordingly, the updated frame marker 425-b is generated by performing windowed self-attention based on global marker 405, text marker 415, each marker within the current local window 430, previous keyframes, and subsequent keyframes in subsequent windows (such as subsequent window 435).

[0079] like Figure 4 As shown, each local window includes a keyframe and a set of four frame markers; for example, local window 430 includes keyframe 420 and frame marker 425. However, Figure 4 This is merely an indication. The embodiments are not limited to this; in some examples, a window may include approximately 2000 markers. Accordingly, a frame marker refers to multiple unique markers, such as... Figure 4 As shown, each marker ( Figure 4 Each bar depicting frame marker 425 can represent approximately 500 unique markers. According to one exemplary embodiment, a global marker (such as global marker 405) can refer to a size of 512. A vector of 2048. In some examples, such as... Figure 4 As shown, the text tag can refer to a size of 3 A vector of 2048.

[0080] One embodiment of this disclosure is configured to vary the window size at different attention layers of the diffusion network. In some cases, the video generation model uses a first window size (such as the window size of local window 430) at a first attention layer of the diffusion network. In other cases, the video generation model uses a second window size different from the first window size at a second attention layer of the diffusion network, which is different from the first attention layer.

[0081] One embodiment of this disclosure is configured to perform a window movement operation. In some cases, the video generation model moves the position of a window across attention layers, such as from a local window 430 associated with frame marker 425-a to a subsequent window 435 associated with a subsequent marker. For example, the video generation model performs position movement to enable the propagation of global information across the attention layers of the diffusion network. In some examples, by performing window position movement, embodiments of this disclosure can improve the efficiency of global information propagation.

[0082] Figure 5 An example of a transformer network 500 according to various aspects of this disclosure is shown. The example includes a transformer 500, an encoder 505, a decoder 520, an input 540, an input embedding 545, an input position encoding 550, a previous output 555, a previous output embedding 560, a previous output position encoding 565, and an output 570. According to some aspects, the encoder 505 is implemented as a video encoder of a multimodal encoder. According to some aspects, the encoder 505 is implemented as a text encoder of a conditional text encoder. According to some aspects, the transformer 500 is used in video generation models (such as reference...) Figure 4 , Figures 6 to 9 and Figure 17 The video generation model described is implemented in this way.

[0083] In some cases, encoder 505 includes a multi-head self-attention sublayer 510 and a feedforward network sublayer 515. In some cases, decoder 520 includes a first multi-head self-attention sublayer 525, a second multi-head self-attention sublayer 530, and a feedforward network sublayer 535.

[0084] In some cases, encoder 505 is configured to map input 540 (e.g., a text prompt) to a sequence of continuous representations fed into decoder 520. In some cases, decoder 520 generates output 570 (e.g., a prediction of the output sequence of words or tokens) based on the output of encoder 505 and previous output 555 (e.g., a previously predicted output sequence), which allows the use of autoregression.

[0085] For example, in some cases, encoder 505 parses input 540 into tokens, vectorizes the parsed tokens to obtain input embedding 545, and adds input position encoding 550 (e.g., a position encoding vector of input 540 with the same dimension as input embedding 545) to input embedding 545. In some cases, input position encoding 550 includes information about the relative positions of words or tokens in input 540.

[0086] In some cases, encoder 505 includes one or more encoding layers that generate contextualized token representations, where each representation corresponds to a token that combines information from other input tokens via a self-attention mechanism. In some cases, each encoding layer of encoder 505 includes a multi-head self-attention sublayer (e.g., multi-head self-attention sublayer 510). In some cases, the multi-head self-attention sublayer implements a multi-head self-attention mechanism that receives different linear projection versions of the query, key, and value to produce outputs in parallel. In some cases, each encoding layer of encoder 505 also includes a fully connected feedforward network sublayer (e.g., feedforward network sublayer 515) that includes two linear transformations around the activation of a rectified linear unit (ReLU):

[0087] In some cases, different weight parameters are used for each layer. and different bias parameters Apply the same linear transformation to each word or token in the input 540.

[0088] In some cases, each sub-layer of encoder 505 is followed by a normalization layer, which normalizes the input at the sub-layer. x With the output sublayer generated by this sublayer ( x Normalize the sum calculated between )

[0089] In some cases, encoder 505 is bidirectional because encoder 505 applies attention to each word or token in input 540, regardless of the word or token's position in input 540.

[0090] According to some aspects, encoder 505 serves as a text encoder for a conditional text encoder. In one example, the conditional text encoder splits the input text into fixed-size segments, generates a linear embedding for each segment, adds a positional embedding to each linear embedding, and provides the resulting vector sequence as input 540 to encoder 505.

[0091] In some cases, decoder 520 includes one or more decoding layers (e.g., six decoding layers). In some cases, each decoding layer includes three sub-layers, including a first multi-head self-attention sub-layer (e.g., first multi-head self-attention sub-layer 525), a second multi-head self-attention sub-layer (e.g., second multi-head self-attention sub-layer 530), and a feedforward network sub-layer (e.g., feedforward network sub-layer 535). In some cases, each sub-layer of decoder 520 is followed by a normalization layer that normalizes the input at the sub-layer.x With the output sublayer generated by this sublayer ( x The sum calculated between ) is normalized.

[0092] In some cases, decoder 520 generates a previous output embedding 560 of the previous output 555 and adds a previous output position code 565 (e.g., position information of words or tokens in the previous output 555) to the previous output embedding 560. In some cases, each first multi-head self-attention sub-layer receives a combination of the previous output embedding 560 and the previous output position code 565 and applies a multi-head self-attention mechanism to this combination. In some cases, for each word in the input sequence, each first multi-head self-attention sub-layer of decoder 520 only applies attention to the words preceding that word in the sequence, so the transformer 500's prediction of the word at a particular position depends only on the known outputs of the words preceding that word in the sequence. For example, in some cases, each first multi-head self-attention sub-layer implements multiple single attention functions in parallel by introducing a mask on the values ​​produced by the scaling multiplication of matrices Q and K, and by suppressing matrix values ​​that would otherwise correspond to disallowed connections.

[0093] In some cases, by receiving a query Q from the previous sublayer of decoder 520 and a key K and a value V from the output of encoder 505, each second multi-head self-attention sublayer implements a multi-head self-attention mechanism similar to that implemented in each multi-head self-attention sublayer of encoder 505, thereby allowing decoder 520 to apply attention to each word in input 540.

[0094] In some cases, each feedforward sublayer implements a fully connected feedforward network similar to feedforward sublayer 515. In other cases, the feedforward sublayer is followed by a linear transformation and softmax to generate the prediction of output 570.

[0095] Figure 6 Examples of guided diffusion models 600 according to various aspects of this disclosure are shown. In some examples, the guided diffusion model 600 describes a reference... Figure 17 The operation and architecture of the video generation model 1715 are described. Figure 6 The guided potential diffusion model 600 described herein is an example of the video generation model described herein, or includes aspects thereof.

[0096] Diffusion models are a class of generative neural networks that can be trained to generate new data with features similar to those found in the training data. Specifically, diffusion models can be used to generate novel media items such as images, audio files, videos, 3D models, or other digital media items. Diffusion models can be used for a variety of media processing tasks, including image super-resolution, generation of media items with perceptual metrics, conditional generation (e.g., text-guided generation), image inpainting, and other media processing tasks.

[0097] The diffusion model works by iteratively adding noise to the data during the forward process and then learning to recover the data by denoising it during the backward process. For example, during training, the guided latent diffusion model 600 can take the original media item 605 in the pixel space 610 as input and apply the forward diffusion process 615 to gradually add noise to the original media item 605 to obtain noisy media items 620 at various noise levels.

[0098] Next, the backdiffusion process 625 (e.g., U-Net) progressively removes noise from the noisy media term 620 at various noise levels to obtain the output media term 630. In some cases, the output media term 630 is created from each of the various noise levels. The output media term 630 can be compared with the original media term 605 to train the backdiffusion process 625.

[0099] The backdiffusion process 625 can also be guided based on text cues 635 or another guiding cue (such as an image, layout, segmentation map, etc.). Text cues 635 can be encoded using a text encoder 665 (e.g., a multimodal encoder) to obtain guiding features 645 in the guiding space 650. Guiding features 645 can be combined with noisy media items 620 at one or more layers of the backdiffusion process 625 to ensure that the output media item 630 includes the content described by the text cues 635. For example, guiding features 645 can be combined with noisy features using cross-attention blocks within the backdiffusion process 625. For example, the backdiffusion process 625 can be a diffusion transformer network (such as a reference...) Figure 7 The described diffusion converter network) or U-Net (such as reference) Figure 8 The U-Net 800 described.

[0100] Methods for manipulating diffusion models include the Denoising Diffusion Probabilistic Model (DDPM) and the Denoising Diffusion Implicit Model (DDIM). In DDPM, the generative process involves a reverse stochastic Markov diffusion process. DDIM, on the other hand, uses a deterministic process that ensures the same input produces the same output. In some cases, DDIM can reduce the number of time steps during media generation. Diffusion models can also be characterized by whether noise is added to the media item itself or to the media features generated by the encoder (i.e., latent diffusion). In pixel-based diffusion models, noise is added and removed in the pixel space. In latent diffusion models, noise is added (and removed) in the latent space of the media features rather than in the pixel space. Therefore, latent diffusion models use back-diffusion to generate media features, and these media features can be decoded to obtain the synthetic media item. DDIM is a reference... Figure 2 , Figure 5 , Figures 7 to 9 and Figures 11 to 17 Examples of the corresponding elements described, or including aspects thereof.

[0101] Figure 7 Examples of diffusion transformer (DiT) architectures according to various aspects of this disclosure are shown. The examples shown include a noisy latent 700, a pachify operation 705, a time-step embedding 710, (multiple) DiT blocks, layer normalization 720, linear and reshaping layers 725, prediction noise 730, input labeling 735, conditional input labeling 740, self-attention 745, cross-attention 750, and a feedforward network 755.

[0102] The block division operation 705 is a reference. Figure 4 Examples of the corresponding elements described, or including aspects thereof. Input marker 735 is a reference. Figure 4 and Figure 10 Examples of the corresponding elements described, or including aspects thereof.

[0103] The DiT architecture processes a noisy latent 700, which can be a noisy version of the input image encoded in the latent space. A block division operation 705 divides the noisy latent into a sequence of tiles to be processed as labels. Each label is a vector representation of each tile in the latent space and is adjusted through an attention process. Each label also receives a temporal embedding 710 encoded for the current denoising time step, and a positional embedding encoded for the spatial location of each label in the image. The label and temporal step information are processed through N DiT blocks 715, where N is the number of DiT blocks.

[0104] Each DiT block 715 comprises multiple processing stages. Initially, a pruning operation is performed, where the router model determines which tags to process or skip based on the learned layer-time step adaptive compression ratio. The remaining tags are processed as input tags 735, which interact with the conditional input tags 740 through multiple attention mechanisms. Self-attention 745 allows input tags to pay attention to each other, while cross-attention 750 enables input tags to pay attention to the conditional input tags 740. The output is then processed through a feedforward network 755. This process is repeated for each DiT block in the sequence.

[0105] After processing all DiT blocks, the output undergoes layer normalization 720, followed by linear and reshaping layers 725. The final output is prediction noise 730, which represents the model's prediction of the noise added to create the initial noisy latent 700. Prediction noise 730 is used to remove the noisy latent 700 at each diffusion time step. At the end of the denoising schedule, the latent samples are decoded to generate a synthetic image in pixel space.

[0106] Figure 8 Examples of diffusion transformer networks for video 800 according to various aspects of this disclosure are shown. In some examples, the diffusion transformer network 800 includes a variational autoencoder 810, a diffusion process 820, a visual transformer 830, a transformer encoder (DiT) 850, a linear decoder and reshaping 860, and a decoder 870. In some aspects, the diffusion transformer network 800 generates a visual image block 815, a diffused visual image block 825, a latent vector 835, a denoised latent code 855, a denoised image block 865, and a media item 875.

[0107] As shown in the figure, a training video 805 comprising multiple frames (e.g., 1920×1080 resolution) is processed by a visual encoder 810 (such as a variational autoencoder (VAE) encoder module). The visual encoder 810 generates a latent representation set 815 (referred to herein as visual blocks), which corresponds to a compressed representation of the visual content of the input video 805.

[0108] These visual blocks 815 undergo a diffusion process 820 (such as reference). Figures 6 to 7 and Figure 9 The diffusion process described herein (820) incrementally destroys tiles by adding noise according to a predetermined noise schedule. The diffused visual tiles are then processed by a vision transformer (ViT) module 830 configured to apply both tile-based encoding and positional encoding, thereby producing a latent representation 835 that captures spatial and temporal correlations across video frames.

[0109] In some cases, the diffusion transformer network 800 takes conditional input 845 (including text descriptions, images, and / or videos) as input. The input (latent representation 835, conditional input 845) can be processed using language and visual encoders to generate conditional latent 850. Conditional latent 850 can represent features aligned with the user's intent or semantic constraints.

[0110] Converter encoder 855 (such as diffuse converter (such as reference) Figure 7 The described DiT receives latent representations and conditional latents 850 and processes combined information across a series of transformer layers. The output 860 of the transformer encoder 855 is decoded via a linear decoder, followed by a reshaping operation 865 to generate a sequence of latent representations 870 corresponding to the predicted visual content.

[0111] In some cases, the VAE decoder 875 reconstructs the video 880 from the decoded latent representation 870. The resulting output video 880 can be generated in response to textual or visual cues, or it can reconstruct the content of the original training video.

[0112] Figure 9 A diffusion process 900 according to various aspects of this disclosure is illustrated. In some examples, the diffusion process 900 is described with reference to... Figure 17 The operation of the video generation model 1715 described, such as reference Figure 6 The reverse diffusion process 625 of the guided diffusion model 600 is described.

[0113] As per the above reference Figure 6 The described diffusion model can involve both a forward diffusion process 905 for adding noise to a media item (or feature in the latent space) and a backward diffusion process 910 for denoising the media item (or feature) to obtain a denoised media item. The forward diffusion process 905 can be represented as... The reverse diffusion process 910 can be represented as In some cases, the forward diffusion process 905 is used during training to generate media terms with successively larger noise, and the neural network is trained to perform the backward diffusion process 910 (i.e., to successively remove noise).

[0114] In the example forward pass of the latent diffusion model, the model uses a Markov chain to store the observed variables (in the pixel space or the latent space). Mapping to intermediate variables ... As latent variables are passed through neural networks (such as U-Net), the Markov chain gradually adds Gaussian noise to the data to obtain approximate posterior probabilities. ,in ... With Same dimension.

[0115] Neural networks can be trained to perform a backward diffusion process. During the backward diffusion process 910, the model diffuses data from noisy data. (Starting with noisy media item 915), and denoising the data to obtain... In each step The reverse diffusion process 910 will (such as the first intermediate media item 920) and As input. Here, This represents the steps in the transition sequence associated with different noise levels. The backdiffusion process 910 iteratively outputs... (such as the second intermediate media item 925), until Reply (Original media item 930). The reverse process can be represented as:

[0116] The joint probability of a sequence of samples in a Markov chain can be written as the product of the conditional probability and the marginal probability: in It refers to the pure noise distribution when the reverse process takes the result of the forward process (pure noise samples) as input, and This represents the Gaussian transition sequence corresponding to the Gaussian noise sequence added to the sample.

[0117] Data observed in pixel space during interference It can be mapped into the latent space as input, and the generated data It is mapped back from the latent space to the pixel space as output. In some examples, This represents a latent variable indicating a low-quality original input medium. ... This indicates a noisy media item, while This indicates that the generated items are of high quality.

[0118] Accordingly, an apparatus for video processing is described. One or more aspects of the apparatus include: obtaining an input cue describing a scene; generating a global marker based on the input cue using a video generation model by performing an attention process based on multiple frame markers; generating frame markers using the video generation model by performing an attention process based on the global markers and a subset of multiple frame markers in a local window; and generating a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers. In some aspects, the video generation model includes a diffusion network.

[0119] Some examples of the apparatus and system also include a text encoder configured to generate text tags based on input prompts, wherein global tags and frame tags are generated by performing an attention process based on the text tags. Some examples of the apparatus and system also include generating frame tags by performing an attention process based on multiple keyframes corresponding to multiple frame blocks, each of the multiple frame blocks comprising a unique subset of multiple frame tags.

[0120] Some examples of the apparatus and system also include generating frame markers, which includes: initializing multiple noisy frame markers; and denoising the multiple noisy frame markers based on a global marker to obtain multiple frame markers. Some examples of the apparatus and system also include generating synthetic video, which includes generating multiple image frames corresponding to the multiple frame markers. Video generation process

[0121] This disclosure describes systems and methods for video generation. Embodiments of this disclosure are configured to perform a windowed self-attention mechanism in a video generation model. In some cases, the video generation model of this disclosure is based on a diffusion network. In other cases, the video generation model is configured to perform segmentation of an input video sequence to generate one or more non-overlapping temporal videos.

[0122] Embodiments of this disclosure are configured to perform a non-overlapping windowed self-attention mechanism. In some cases, the video generation model modifies the window size. For example, the video generation model reduces the complexity associated with the window size based on a linear variation. In some examples, the video generation model performs segmentation of the input video sequence based on a variable window size. By changing the window size at different attention layers, embodiments of this disclosure can optimize the efficiency and performance of video generation.

[0123] One embodiment of this disclosure is configured to perform video generation. In some cases, a video generation model is used to learn global tags and global context. For example, the video generation model learns global tags from keyframes associated with each of one or more non-overlapping time videos. For example, the video generation model learns global context from keyframes associated with each of one or more non-overlapping time videos. By using a video generation model to learn global tags and global context from keyframes, embodiments of this disclosure can efficiently propagate the global context used for video generation.

[0124] In some cases, the video generation model is used to implement context from nearby keyframes. For example, nearby keyframes refer to keyframes adjacent to the current window. In other cases, the video generation model is used to implement context from a moving window that moves across attention layers. For example, the window moves across the attention layers of a video generation model that includes a diffusion network. By combining context from nearby keyframes and moving windows, embodiments of this disclosure ensure the consistency and efficient propagation of global context across the diffusion network architecture of the video generation model.

[0125] Figure 10 Examples of a method 1000 for video processing according to aspects of this disclosure are shown. In some examples, these operations are performed by a system including a processor that executes a set of code to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed based on the methods and processes described according to aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.

[0126] In step 1005, the system receives an input prompt. In some cases, this step involves referring to... Figure 1 The user interface described, or that can be executed by it.

[0127] For example, in some cases, video processing devices (such as reference devices) Figure 17 The user interface of the described video processing apparatus 1700 receives input prompts from a user. In some examples, the input prompts include text prompts describing objects that the user wants to depict in the generated asset (e.g., video). Additionally or alternatively, for example, the input prompts include text prompts describing actions performed by elements that the user wants to depict in the generated asset (e.g., video). In some examples, the video processing apparatus receives input prompts from a database or any other data source.

[0128] In operation 1010, the system generates global markers based on input prompts and a set of global markers corresponding to the set of video frames. In some cases, this step involves referencing... Figures 4 to 9 and Figure 17 The described video generation model, or the model that can be executed by it.

[0129] This disclosure describes systems and methods for video generation. Embodiments of this disclosure include video generation models configured to perform windowed self-attention operations. In some cases, the attention operation includes updating global tags (such as reference tags). Figure 4 The process of global tagging (405) and updating frame tags is described. For example, global tags consist of learnable register tags, image tags, video tags, and text tags.

[0130] In some cases, the global label is updated by performing an attention process based on each label to generate an updated global label. For example... Figure 4 As shown in step 1, the global marker 405-a is updated by performing an attention process based on each of the text markers (such as text marker 415), keyframes (such as keyframe 420), and frame markers (such as frame marker 425) to generate an updated global marker 405-b, as indicated by the arrow.

[0131] In operation 1015, the system generates frame markers based on a global marker and a subset of local markers from the global marker set, where the subset of local markers corresponds to a local window that includes a subset of video frames from the video frame set, and where the frame marker corresponds to a frame within the subset of video frames. In some cases, this step involves referencing... Figures 4 to 9 and Figure 17 The described video generation model, or the model that can be executed by it.

[0132] In some cases, in addition to the process of updating the global label described in Operation 1010, the windowed self-attention operation also includes the processes of updating keyframes and updating frame labels. In some cases, the video generation model utilizes local windows (such as reference windows). Figure 4 Step 2 describes updating the tags using a local window 430. For example, the video generation model performs the update by performing an attention process based on each tag within the local window 430. In some cases, frame tags apply attention to global tags.

[0133] Additionally, in some cases, frame tags apply attention to nearby keyframes. For example, frame tags apply attention to previous keyframes and subsequent keyframes of subsequent local tag subsets. See reference... Figure 4As shown in step 2, each frame tag in the local tag subset (i.e., local window 430) applies attention to the previous keyframe and subsequent keyframes associated with subsequent local tag subsets (such as subsequent keyframes associated with subsequent local window 435).

[0134] In operation 1020, the system generates video based on frame tags, where the video includes an image corresponding to each video frame in the set of video frames. In some cases, this step involves referencing... Figures 4 to 9 and Figure 17 The described video generation model, or one that can be executed by it. In some embodiments, the user interface provides the video to the user.

[0135] One embodiment of this disclosure is configured to change the window size. In some cases, the video generation model uses a first window size at a first attention layer of the diffusion network. In other cases, the video generation model uses a second window size different from the first window size at a second attention layer of the diffusion network, which is different from the first attention layer.

[0136] One embodiment of this disclosure is configured to perform a window movement operation. In some cases, the video generation model moves the window's position across attention layers. For example, the video generation model performs position movement to enable the propagation of global information across the attention layers of the diffusion network. In some examples, by performing window position movement, embodiments of this disclosure can improve the efficiency of global information propagation.

[0137] Figure 11 Examples of a method 1100 for video processing according to aspects of this disclosure are shown. In some examples, these operations are performed by a system including a processor that executes a set of code to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed based on the methods and processes described according to aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.

[0138] In step 1105, the system receives an input prompt. In some cases, this step involves referring to... Figure 17 The video processing apparatus described, or that can be performed by it.

[0139] For example, in some cases, video processing devices (such as reference devices) Figure 17The user interface of the described video processing apparatus 1700 receives input prompts from a user. In some examples, the input prompts include text prompts describing objects that the user wants to depict in the generated asset (e.g., video). Additionally or alternatively, for example, the input prompts include text prompts describing actions performed by elements that the user wants to depict in the generated asset (e.g., video). In some examples, the video processing apparatus receives input prompts from a database or any other data source.

[0140] In operation 1110, the system generates frame markers based on input prompts and a subset of local markers from a global marker set corresponding to the set of video frames, wherein the subset of local markers corresponds to a local window that includes a subset of video frames from the set of video frames, and wherein the frame markers correspond to frames within the subset of video frames. In some cases, this step involves referencing... Figure 17 The described video generation model, or the model that can be executed by it.

[0141] In operation 1115, the system generates an updated frame tag based on the frame tag and different local tag subsets from the global tag set, where the different local tag subsets correspond to different local windows with window sizes different from the local windows. In some cases, this step involves referencing... Figure 17 The described video generation model, or the model that can be executed by it.

[0142] In some cases, attention operations involve updating frame tags simultaneously with updating global and text tags. In other cases, the video generation model updates frame tags within a local window (such as updating frame tag 425-a in local window 430 to generate updated frame tag 425-b, as referenced). Figure 4 (as described in step 2). For example, the video generation model performs updates by executing an attention process based on each frame tag within a local window. In some cases, the frame tags are relative to the global tag and text tags (such as...). Figure 4 Attention is applied to the global marker 405 and the text marker 415 in step 2.

[0143] Additionally, in some cases, frame tags apply attention to nearby keyframes. For example, frame tags apply attention to previous keyframes and subsequent keyframes of subsequent local windows. See reference... Figure 4 As shown in step 2, each frame marker 425 in the local window 430 applies attention to the previous keyframe and subsequent keyframes associated with subsequent local marker subsets (such as subsequent keyframes associated with subsequent local windows 435).

[0144] In operation 1120, the system generates video based on the updated frame tags, where the video includes an image corresponding to each video frame in the video frame set. In some cases, this step involves referencing... Figure 17 The described video generation model, or one that can be executed by it. In some embodiments, the user interface provides the video to the user.

[0145] Figure 12 Examples of a method 1200 for video processing according to aspects of this disclosure are shown. In some examples, these operations are performed by a system including a processor that executes a set of code to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed based on the methods and processes described according to aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.

[0146] In step 1205, the system receives input prompts describing the scenario. In some cases, this step involves referencing... Figure 17 The video processing apparatus described, or that can be performed by it.

[0147] For example, in some cases, video processing devices (such as reference devices) Figure 17 The user interface of the described video processing apparatus 1700 receives input prompts from a user. In some examples, the input prompts include text prompts describing a scene that the user wants to generate an asset (e.g., a composite video) based on. Additionally or alternatively, for example, the input prompts include text prompts describing actions performed by elements in the scene that the user wants to depict in the composite video. In some examples, the video processing apparatus receives input prompts from a database or any other data source.

[0148] In operation 1210, the system generates global tags based on input cues by performing an attention process based on a set of frame tags. In some cases, this step involves referencing... Figure 17 The described video generation model, or the model that can be executed by it.

[0149] One embodiment of this disclosure includes a video generation model configured to perform windowed self-attention operations. In some cases, the attention operation includes a process of updating global labels (such as updating global label 405-a to generate an updated global label 405-b, as referenced). Figure 4(as described in step 1). For example, global tags consist of learnable register tags, image tags, video tags, and text tags. In some cases, global tags are based on multiple tags (such as global tag 405-a indicated by the arrow, text tag 415-a, and multiple tags 410, as shown in reference). Figure 4 Each label in step 1 (as described) is updated using self-attention. Figure 4 As shown, the multiple markers 410 include keyframe 420 and frame marker 425.

[0150] In operation 1215, the system generates frame tags by performing an attention process based on a subset of the set of frame tags in the global tag and the local window. In some cases, this step involves referencing... Figure 17 The described video generation model, or the model that can be executed by it.

[0151] In some cases, the video generation model generates frame markers based on global markers and frame markers within a local window, where the frame markers correspond to a local window that includes a subset of video frames from the set of video frames, and where the frame markers correspond to frames within the subset of video frames.

[0152] In some cases, attention operations involve updating frame tags simultaneously with updating global and text tags. In some cases, such as... Figure 4 The video generation model updates the frame markers within a local window (e.g., updating frame marker 425-a in local window 430 to generate an updated frame marker 425-b, as referenced). Figure 4 (as described in step 2). Additionally, for example, the video generation model updates frame marker 425-a by performing a windowed self-attention operation based on each frame marker within a local window 430. In some cases, frame marker 425-a applies attention to global marker 405, text marker 415, and keyframe 420.

[0153] Additionally, in some cases, frame tags apply attention to nearby keyframes. For example, frame tags apply attention to previous keyframes and subsequent keyframes of subsequent local tag subsets. See reference... Figure 4 As shown in step 2, each frame marker 425-a in local window 430 applies attention to previous keyframes and subsequent keyframes associated with subsequent local windows (such as subsequent keyframes associated with subsequent local window 435) (as indicated by the arrows) to generate an updated frame marker 425-b.

[0154] In operation 1220, the system generates a synthetic video based on frame markers, wherein the synthetic video depicts a scene and includes image frames corresponding to the frame markers. In some cases, this step involves referencing... Figure 17The described video generation model, or one that can be executed by it. In some embodiments, the user interface provides the video to the user.

[0155] Figure 13 Examples of a method 1300 for video processing according to aspects of this disclosure are shown. In some examples, these operations are performed by a system including a processor that executes a set of code to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed based on the methods and processes described according to aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.

[0156] In step 1305, the system receives input prompts describing the scenario. In some cases, this step involves referencing... Figure 17 The video processing apparatus described, or that can be performed by it.

[0157] For example, in some cases, video processing devices (such as reference devices) Figure 17 The user interface of the described video processing apparatus 1700 receives input prompts from a user. In some examples, the input prompts include text prompts describing a scene that the user wants to generate an asset (e.g., a composite video) based on. Additionally or alternatively, for example, the input prompts include text prompts describing actions performed by elements in the scene that the user wants to depict in the composite video. In some examples, the video processing apparatus receives input prompts from a database or any other data source.

[0158] In operation 1310, the system generates text tags based on input prompts. In some cases, this step involves referencing... Figure 17 The text encoder described, or that can be executed by it.

[0159] According to one embodiment, a text encoder is configured to process input prompts and generate text tokens. For example, input prompts refer to text input including natural language phrases, captions, sentences, other structured or unstructured language content, etc. In some examples, the text encoder may include one or more neural network architectures (such as transformer-based models, recurrent neural networks, or convolutional networks) that receive input prompts and perform tokenization, embedding, and context encoding operations to generate corresponding text tokens. For example, the generated text tokens may be represented as high-dimensional vectors capturing the semantic and syntactic properties of the input prompts.

[0160] In operation 1315, the system generates frame tags by performing an attention process based on a subset of the set of text tags and frame tags in a local window. In some cases, this step involves referencing... Figure 17 The described video generation model, or the model that can be executed by it.

[0161] In some cases, the video generation model generates frame markers based on global markers, text markers, and frame markers, where a frame marker corresponds to a local window that includes a subset of video frames from a set of video frames, and where a frame marker corresponds to a frame within the subset of video frames.

[0162] In some cases, windowed self-attention operations update frame tags, where the self-attention operation includes updating frame tags and simultaneously updating global and text tags. In other cases, the video generation model updates frame tags within a local window (e.g., updating frame tag 425-a in local window 430 to generate updated frame tag 425-b, as shown in reference...). Figure 4 (as described in step 2). For example, the video generation model performs updates by executing an attention process based on each frame tag within a local window. In some cases, attention is applied to the global tag, the text tag, and each frame tag within the local window.

[0163] Additionally, in some cases, frame tags apply attention to nearby keyframes. For example, frame tags apply attention to previous keyframes and subsequent keyframes of subsequent local tag subsets. See reference... Figure 4 As shown in step 2, each frame tag in the local tag subset 430 applies attention to the previous keyframe and subsequent keyframes associated with subsequent local tag subsets (such as subsequent keyframes associated with subsequent local windows 435).

[0164] In operation 1320, the system generates a synthetic video based on frame markers, wherein the synthetic video depicts a scene and includes image frames corresponding to the frame markers. In some cases, this step involves referencing... Figure 17 The described video generation model, or one that can be executed by it. In some embodiments, the user interface provides the video to the user.

[0165] One embodiment of this disclosure is configured to change the window size. In some cases, the video generation model uses a first window size at a first attention layer of the diffusion network. In other cases, the video generation model uses a second window size different from the first window size at a second attention layer of the diffusion network, which is different from the first attention layer.

[0166] One embodiment of this disclosure is configured to perform a window movement operation. In some cases, the video generation model moves the window's position across attention layers. For example, the video generation model performs position movement to enable the propagation of global information across the attention layers of the diffusion network. In some examples, by performing window position movement, embodiments of this disclosure can improve the efficiency of global information propagation.

[0167] Accordingly, a method for video processing is described. One or more aspects of the method include: obtaining an input cue describing a scene; generating a global marker based on the input cue by performing an attention process based on multiple frame markers using a video generation model; generating frame markers by performing an attention process based on the global markers and a subset of multiple frame markers in a local window using the video generation model; and generating a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0168] Examples of methods, apparatuses, and non-transitory computer-readable media also include generating text tags based on input prompts, where global tags and frame tags are generated by performing an attention process based on text tags.

[0169] Some examples of the method, apparatus, and non-transitory computer-readable medium also include generating frame markers, which involves performing an attention process based on multiple keyframes that correspond to multiple frame blocks, each of the multiple frame blocks comprising a unique subset of the multiple frame markers.

[0170] Examples of the method, apparatus, and non-transitory computer-readable medium include generating frame markers, which includes: initializing a plurality of noisy frame markers; and denoising the plurality of noisy frame markers based on a global marker to obtain a plurality of frame markers.

[0171] Some examples of the method, apparatus, and non-transitory computer-readable medium also include iteratively updating a global tag based on a frame tag. Some examples also include iteratively updating a frame tag based on a global tag.

[0172] Some examples of the method, apparatus, and non-transitory computer-readable medium also include generating a synthetic video, which includes generating multiple image frames corresponding to multiple frame markers.

[0173] Some examples of the method, apparatus, and non-transitory computer-readable medium also include identifying the diffusion time step. Some examples also include determining whether to perform selective attention on frame markers based on the diffusion time step, wherein the selective attention is based on a local window.

[0174] Additionally, a method for video processing is described. One or more aspects of this method include: obtaining input cues describing a scene; generating text markers based on the input cues; generating frame markers using a video generation model by performing an attention process based on the text markers and a subset of frame markers in a local window; and generating a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0175] Examples of methods, apparatuses, and non-transitory computer-readable media also include: using a video generation model to generate global markers based on input cues by performing an attention process based on multiple frame markers, wherein the frame markers are based on global markers.

[0176] Examples of the method, apparatus, and non-transitory computer-readable medium include generating frame markers, which includes: initializing a plurality of noisy frame markers; and denoising the plurality of noisy frame markers based on a global marker to obtain a plurality of frame markers.

[0177] Some examples of the method, apparatus, and non-transitory computer-readable medium also include iteratively updating a global tag based on a frame tag. Some examples also include iteratively updating a frame tag based on a global tag.

[0178] Some examples of the method, apparatus, and non-transitory computer-readable medium also include generating frame markers, which involves performing an attention process based on multiple keyframes that correspond to multiple frame blocks, each of the multiple frame blocks comprising a unique subset of the multiple frame markers.

[0179] Some examples of the method, apparatus, and non-transitory computer-readable medium also include generating a synthetic video, which includes generating multiple image frames corresponding to multiple frame markers.

[0180] Some examples of the method, apparatus, and non-transitory computer-readable medium also include identifying the diffusion time step. Some examples also include determining whether to perform selective attention on frame markers based on the diffusion time step, wherein the selective attention is based on a local window. Training System

[0181] One embodiment of this disclosure is configured to perform video generation. In some cases, the video generation model is used to learn global tags and global context. For example, the video generation model learns global tags from keyframes associated with each of one or more non-overlapping temporal videos.

[0182] In some cases, a global tag refers to a compact representation of the input video sequence. For example, a global tag captures long-range correlations based on the input video sequence. Accordingly, by combining learnable global tags from keyframes associated with each of one or more non-overlapping temporal videos, embodiments of this disclosure are able to capture long-range correlations while reducing the computational cost associated with self-attention mechanisms.

[0183] One embodiment of this disclosure is configured to implement learnable register tagging during the video generation process. In some cases, the video generation model uses learnable register tagging to propagate global context between windows at each attention layer of the diffusion network. In other cases, the video generation model is configured to generate video based on a diffusion network architecture.

[0184] Figure 14 Examples of methods for training machine learning models according to various aspects of this disclosure are shown. Figure 14 This is a flowchart of a step-by-step process 1400 in an example implementation of an algorithm described as an executable operation for training a machine learning model. In some embodiments, process 1400 describes the training component 1730 as being configured with a reference. Figure 17 The operation of the video generation model 1715 is described. Process 1400 provides one or more examples of generating training data, using the training data to train a machine learning model, and using the trained machine learning model to perform a task.

[0185] In this example, it begins with the collection of training data from the machine learning system (Box 1402). This training data will be used as the basis for training the machine learning model; that is, the training data defines what is being modeled. Training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, service provider system platforms with publicly available application programming interfaces (e.g., social media platforms), user data collection systems (e.g., digital surveys and online crowdsourcing systems), and so on. Training data collection may also include: data augmentation and synthetic data generation techniques to expand and diversify the available training data; balancing techniques to balance multiple positive and negative examples; and so on.

[0186] The machine learning system can also be configured to identify features relevant to a task type (box 1404), and the machine learning model will be trained for that task type. Examples of tasks include classification, natural language processing, generative artificial intelligence, recommendation engines, reinforcement learning, clustering, and so on. To this end, the machine learning system collects training data based on the identified features, and / or filters the training data based on the identified features after collection. The training data is then used to train the machine learning model.

[0187] To train the machine learning model in the example shown, the machine learning model is first initialized (box 1406). Initializing the machine learning model includes selecting the model architecture to train (box 1408). Examples of model architectures include neural networks, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning neural networks, etc.

[0188] Also select a loss function (box 1410). The loss function is used to measure the difference between the output (i.e., prediction) of the machine learning model and the target value (e.g., represented by the training data) to be used to train the machine learning model. Additionally, select (1412) an optimization algorithm that will be used in conjunction with the loss function to optimize the parameters of the machine learning model during training. Examples of optimization algorithms include gradient descent, stochastic gradient descent (SGD), and so on.

[0189] Initializing a machine learning model also includes setting initial values ​​for the model (box 1414). Examples of this initialization include initializing node weights and biases as part of training to improve the efficiency of training and computational resource consumption. Hyperparameters are also set; these are used to control the training of the machine learning model. Examples of hyperparameters include regularization parameters, model parameters (e.g., the number of layers in a neural network), learning rate, batch size selected from training data, etc. Hyperparameters are set using various techniques, including randomization, heuristics learned from other training scenarios, and so on.

[0190] The machine learning system then uses training data to train a machine learning model (Box 1418). A machine learning model is a computer representation that can be tuned (e.g., trained and retrained) based on inputs of training data to approximate an unknown function. Specifically, the term machine learning model can include models that utilize algorithms (e.g., using the model architecture described above) to learn from known data by analyzing training data and to make predictions on known data, in order to learn and relearn to generate outputs that reflect the patterns and properties described by the training data.

[0191] Examples of training types include supervised learning using labeled data, unsupervised learning involving finding underlying structures or patterns within the training data, reinforcement learning based on optimization functions (e.g., rewards and / or penalties), using nodes as part of “deep learning,” and so on. For example, a machine learning model can be configured to include multiple nodes that collectively form multiple layers. These layers can be configured to include an input layer, an output layer, and one or more hidden layers. Computation is performed by the nodes within these layers through the hidden states via a system of weighted connections “learned” during training (e.g., by using a chosen loss function and backpropagation to optimize the performance of the machine learning model in performing a relevant task).

[0192] As part of training the machine learning model, a determination is made regarding whether stopping criteria are met (decision box 1420), i.e., this determination is used to validate the machine learning model. Stopping criteria can be used to reduce overfitting of the machine learning model, reduce computational resource consumption, and improve the machine learning model's ability to handle previously unseen data (i.e., data not specifically included in the training data as examples). Examples of stopping criteria include, but are not limited to, a predefined number of training epochs, validating loss stabilization, reaching a performance improvement threshold, whether a threshold accuracy level has been met, or based on performance metrics such as precision and recall. If the stopping criteria are not met ("No" from decision box 1420), then in this example, process 1400 continues training the machine learning model using the training data (box 1418).

[0193] If the stopping criterion ("Yes" from decision box 1420) is met, the trained machine learning model is used to generate output based on subsequent data (box 1422). For example, the trained machine learning model is trained to perform the task described above, and therefore, once trained, it is configured to perform the task based on subsequent data received as input and processed by the machine learning model.

[0194] Figure 15 An example of a method for training a diffusion model 1500 according to various aspects of this disclosure is shown. In some embodiments, method 1500 describes a training component 1730 as being configured for reference. Figure 17 The operation of the video generation model 1715 is described. Method 1500 represents the method used for training as shown in the above reference. Figures 7 to 9 Examples of the described reverse diffusion process. In some examples, these operations are performed by a system including a processor that executes a set of code to control the functional elements of the device, such as... Figure 6 The guided diffusion model described in [the text].

[0195] Additionally or alternatively, certain processes of method 1500 may be performed using dedicated hardware. Generally, these operations are performed based on the methods and processes described in accordance with various aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.

[0196] refer to Figure 15 According to some aspects, training components (such as reference) Figure 17 The described training components 1730) train the diffusion model (such as references) Figures 6 to 9 The described video generation model is used to generate output.

[0197] In operation 1505, the user initializes the untrained model. Initialization may include defining the model's architecture and establishing initial values ​​for the model parameters. In some cases, initialization may include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, and the locations of skipped connections.

[0198] In operation 1510, the system usage is divided into... N Each stage of the forward diffusion process (such as reference) Figure 7 The described forward diffusion process adds noise to the training image (or supplementary training images). In some cases, this step involves a reference... Figure 17 The training component described, or that can be executed by it.

[0199] In operation 1515, at each stage From the stage Initially, the system uses a backdiffusion process to predict the phase. The output or features of the input. For example, the backdiffusion process can predict noise added by the forward diffusion process, and the predicted noise can be removed from the noise input to obtain the predicted output. In some cases, the original media terms are predicted at each stage of the training process.

[0200] During operation 1520, the system will enter a phase. The predicted output (or feature) and the actual media item (or feature) (such as stage) The output or original input of the observed data is compared. The diffusion model can be trained to minimize the negative log-likelihood of the training data. The upper limit of the change.

[0201] In operation 1525, the system updates the model's parameters based on this comparison. For example, U-Net's parameters can be updated using gradient descent. Time-dependent parameters of the Gaussian transition can also be learned.

[0202] Accordingly, a method for video processing is described. One or more aspects of the method include: obtaining an input cue; generating frame markers using a video generation model based on the input cue and a subset of local markers from a global marker set corresponding to a set of video frames, wherein the subset of local markers corresponds to a local window comprising a subset of video frames from the set of video frames, and wherein the frame markers correspond to frames within the subset of video frames; generating updated frame markers using the video generation model based on the frame markers and different subsets of local markers from the global marker set, wherein the different subsets of local markers correspond to different local windows having window sizes different from the local windows; and generating a video using the video generation model based on the updated frame markers, wherein the video comprises an image corresponding to each video frame in the set of video frames. One or more aspects of the apparatus include generating global markers based on the input cue and the global marker set, wherein the frame markers are generated based on the global markers. computing devices

[0203] An exemplary embodiment of this disclosure is configured to evaluate the performance of a video generation model. In some cases, the video generation model is configured to perform non-overlapping windowed attention operations while varying the window size at different attention layers of the diffusion network. In other cases, the video generation model is configured to incorporate the context of moving windows from nearby keyframes and attention layers of the diffusion network. For example, the video generation model of this disclosure is capable of significantly improving the speed of video generation (e.g., up to 2x) based on windowed attention operations.

[0204] Figure 16 An example of a computing device according to various aspects of this disclosure is shown. The computing device 1600 may be a reference. Figure 17 An example of the described video processing apparatus 1700. In one aspect, the computing device 1600 includes a plurality of processors 1605, a memory subsystem 1610, a communication interface 1615, an I / O interface 1620, a plurality of user interface components 1625, and a channel 1630.

[0205] In some embodiments, computing device 1600 is Figure 17 Examples of video generation models, or aspects thereof. In some embodiments, computing device 1600 includes one or more processors 1605 that can execute instructions stored in memory subsystem 1610 to perform media generation.

[0206] According to some aspects, computing device 1600 includes one or more processors 1605. In some cases, the processor is an intelligent hardware device, such as a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof. In some cases, the processor is configured to use a memory controller to operate a memory array. In other cases, the memory controller is integrated into the processor. In some cases, the processor is configured to execute computer-readable instructions stored in memory to perform various functions. In some embodiments, the processor includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0207] According to some aspects, the memory subsystem 1610 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software containing instructions that, when executed, cause a processor to perform the various functions described herein. In some cases, the memory contains a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller operates the storage elements. For example, a memory controller may include a row decoder, a column decoder, or both. In some cases, the storage elements within the memory store information in the form of logical states.

[0208] According to some aspects, communication interface 1615 operates at the boundary between communication entities (such as computing device 1600, one or more user devices, a cloud, and one or more databases) and channel 1630, and can record and process communications. In some cases, communication interface 1615 is provided to enable a processing system to couple to a transceiver (e.g., a transmitter and / or a receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna.

[0209] Depending on the context, the I / O interface 1620 is controlled by an I / O controller to manage input and output signals of the computing device 1600. In some cases, the I / O interface 1620 manages peripherals not integrated into the computing device 1600. In some cases, the I / O interface 1620 represents a physical connection or port to an external peripheral. In some cases, the I / O controller uses an operating system such as iOS®, Android®, MS-DOS®, MS-WINDOWS®, OS / 2®, UNIX®, LINUX®, or other known operating systems. In some cases, the I / O controller represents or interacts with a modem, keyboard, mouse, touchscreen, or similar device. In some cases, the I / O controller is implemented as a component of the processor. In some cases, a user interacts with the device via the I / O interface 1620 or via hardware components controlled by the I / O controller.

[0210] According to some aspects, the user interface components 1625 enable a user to interact with the computing device 1600. In some cases, the user interface components 1625 include audio devices (such as external speaker systems), external display devices (such as displays), input devices (e.g., remote control devices that interface directly with the user interface or via an I / O controller), or combinations thereof. In some cases, the user interface components 1625 include a GUI.

[0211] Figure 17 An example of a video processing apparatus 1700 according to various aspects of this disclosure is shown. The video processing apparatus 1700 is a reference. Figure 1 and Figure 3 Examples of the corresponding elements described, or including aspects thereof.

[0212] In one aspect, the video processing apparatus 1700 includes a processor unit 1705, a memory unit 1710, an I / O module 1725, and a training component 1730. In one aspect, the memory unit 1710 includes a video generation model 1715 and a text encoder 1720. The training component 1730 updates the parameters of the video generation model 1715 and the text encoder 1720 stored in the memory unit 1710. In some examples, the training component 1730 is located outside the video processing apparatus 1700. According to some aspects, the video processing apparatus 1700 obtains input cues describing a scene.

[0213] According to some aspects, processor unit 1705 includes processing devices coupled to memory components. Processor unit 1705 includes one or more processors. The processor is an intelligent hardware device, such as a general-purpose processing component, digital signal processor (DSP), central processing unit (CPU), graphics processing unit (GPU), microcontroller, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic device, discrete gate or transistor logic component, discrete hardware component, or any combination thereof.

[0214] In some cases, processor unit 1705 is configured to use a memory controller to operate the memory array. In other cases, the memory controller is integrated into processor unit 1705. In some cases, processor unit 1705 is configured to execute computer-readable instructions stored in memory unit 1710 to perform various functions. In some aspects, processor unit 1705 includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing. According to some aspects, processor unit 1705 includes reference... Figure 16 One or more processors as described.

[0215] Memory unit 1710 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause at least one processor of processor unit 1705 to perform the various functions described herein.

[0216] In some cases, memory cell 1710 includes a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, memory cell 1710 includes a memory controller that operates the storage elements of memory cell 1710. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the storage elements within memory cell 1710 store information in the form of logical states. According to some aspects, memory cell 1710 is a reference... Figure 16 An example of the described memory subsystem 1610.

[0217] According to some aspects, the video processing apparatus 1700 uses one or more processors of the processor unit 1705 to execute instructions stored in the memory unit 1710 to perform the functions described herein. For example, the video processing apparatus 1700 may: obtain an input cue describing a scene; use a video generation model to generate a global marker based on the input cue by performing an attention process based on multiple frame markers; use a video generation model to generate frame markers by performing an attention process based on the global markers and a subset of multiple frame markers in a local window; and use a video generation model to generate a synthetic video based on the frame markers, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0218] In one aspect, memory unit 1710 includes a video generation model 1715, which is trained to: obtain an input cue describing a scene; generate a global marker based on the input cue by performing an attention process based on multiple frame markers using the video generation model; generate frame markers by performing an attention process based on the global markers and a subset of multiple frame markers in a local window using the video generation model; and generate a synthetic video based on the frame markers using the video generation model, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0219] For example, after training, video generation model 1715 can perform reference... Figures 1 to 3 The described inference operations are as follows: obtaining input cues describing a scene; using a video generation model to generate global markers based on the input cues by performing an attention process based on multiple frame markers; using a video generation model to generate frame markers by performing an attention process based on the global markers and a subset of multiple frame markers in a local window; and using a video generation model to generate a synthetic video based on the frame markers, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

[0220] In some embodiments, the video generation model 1715 is an artificial neural network (ANN) comprising multiple networks, including a reference network. Figure 6 The described guided diffusion model and reference Figure 5 The described transformer network. An ANN can be a hardware or software component comprising connection nodes loosely corresponding to neurons in the human brain (i.e., artificial neurons). Each connection or edge transmits a signal from one node to another (like a physical synapse in the brain). When a node receives a signal, it processes it and then transmits the processed signal to other connection nodes.

[0221] ANNs have many parameters, including weights and biases associated with each neuron in the network. These weights and biases control the degree of connection between neurons and affect the neural network's ability to capture complex data patterns. These parameters (also known as model parameters or model weights) are variables that determine the behavior and characteristics of a machine learning model.

[0222] In some cases, the signals between nodes consist of real numbers, and each node's output is computed as a function of its inputs. For example, nodes may use other mathematical algorithms (such as selecting the maximum value from the input as the output) or any other suitable algorithm used to activate the node to determine their output. Each node and edge is associated with one or more node weights, which determine how the signal is processed and transmitted. In some cases, nodes have a threshold below which no signal is transmitted at all. In some examples, nodes are aggregated into layers.

[0223] The parameters of the video generation model 1715 can be organized into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer, and the last layer is called the output layer. In some cases, the signal traverses certain layers multiple times. Hidden (or intermediate) layers contain hidden nodes and are located between the input and output layers. Hidden layers perform nonlinear transformations on the inputs into the network. Each hidden layer is trained to produce a defined output that contributes to the joint output of the ANN's output layers. The hidden representation is a machine-readable representation of the input data learned from the ANN's hidden layers and produced by the output layers. As the understanding of the input ANN improves with training, the hidden representation gradually differs from earlier iterations.

[0224] Training component 1730 can train video generation model 1715 and text encoder 1720. For example, the parameters of video generation model 1715 and text encoder 1720 can be learned or estimated from training data and then used to make predictions or perform tasks based on learned patterns and relationships in the data. In some examples, the parameters are adjusted during training to minimize the loss function or maximize performance metrics (e.g., as referenced). Figure 14 (As described). The goal of the training process can be to discover the optimal parameter values ​​that allow the video generation model to make accurate predictions or perform well on a given task.

[0225] Additionally, in some aspects, the training component 1730 tunes the video generation model 1715 by updating one or more temporal layers. In some cases, the temporal layers generate latent embeddings that influence the motion features of the generated synthetic video. In at least one embodiment, the training component 1730 is implemented on a device different from the video processing device 1700.

[0226] Accordingly, node weights can be adjusted to improve the accuracy of the output (i.e., by minimizing the loss that corresponds in some way to the difference between the current result and the target result). Edge weights increase or decrease the signal strength transmitted between nodes. For example, during training, the algorithm adjusts machine learning parameters according to optimization techniques (such as gradient descent, stochastic gradient descent, or other optimization algorithms) to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the video generation model 1715 can be used to make predictions on new, unseen data (i.e., during inference).

[0227] According to some aspects, the video generation model 1715 generates global markers based on input cues by performing an attention process based on a set of frame markers. In some examples, the video generation model 1715 generates frame markers by performing an attention process based on a global marker and a subset of the set of frame markers in a local window. In some examples, the video generation model 1715 generates synthetic videos based on frame markers, wherein the synthetic videos depict a scene and include image frames corresponding to the frame markers.

[0228] In some examples, the video generation model 1715 generates frame markers by performing an attention process based on a set of keyframes corresponding to a set of frame blocks, wherein each frame block in the set of frame blocks comprises a unique subset of the frame marker set. In some examples, the video generation model 1715 generates frame markers by initializing a noisy set of frame markers and denoising the noisy set of frame markers based on global markers to obtain the frame marker set.

[0229] In some examples, the video generation model 1715 generates synthetic video, including generating a set of image frames corresponding to a set of frame tags. In some examples, the video generation model 1715 identifies diffusion time steps. In some examples, the video generation model 1715 determines whether to perform selective attention on the frame tags based on the diffusion time steps, where selective attention is based on a local window.

[0230] According to some aspects, the video generation model 1715 generates frame tags by performing an attention process based on text tags and a subset of the set of frame tags in a local window. In some examples, the video generation model 1715 generates synthetic videos based on frame tags, wherein the synthetic videos depict scenes and include image frames corresponding to the frame tags. In some aspects, the video generation model 1715 includes a diffusion network.

[0231] Video generation model 1715 is configured to generate synthetic video. Embodiments of video generation model 1715 include image generation models, such as diffusion models adapted to generate temporally coherent frames in a sequence. For example, video generation model 1715 may include a diffusion model with an additional temporal layer.

[0232] According to some aspects, the text encoder 1720 generates text tokens based on input cues, where global tokens and frame tokens are generated by performing an attention process based on the text tokens. In some cases, the text encoder is configured to process text input to generate text tokens. The text input may include natural language phrases, captions, sentences, or other structured or unstructured linguistic content. For example, the text encoder may include one or more neural network architectures, such as transformer-based models (e.g., BERT, GPT, or CLIP), recurrent neural networks (RNNs), or convolutional networks adapted for text. The text encoder receives text input and performs tokenization, embedding, and context encoding operations to generate corresponding text tokens. Text tokens may be represented as high-dimensional vectors capturing the semantic and syntactic properties of the input.

[0233] According to some aspects, a transformer includes an encoder-decoder structure. The encoder of the transformer processes the input sequence and encodes it into a set of high-dimensional representations. The decoder of the transformer generates an output sequence based on the encoded representations and previously generated labels. Both the encoder and decoder include one or more layers of self-attention mechanisms and feedforward ANNs.

[0234] According to one embodiment, a windowed self-attention mechanism can refer to performing self-attention on a set of frames. For example, the windowed attention mechanism is performed on the current window or a local window. In some examples, each local window includes a segmented video sequence containing a set of video frames. In some examples, the windowed attention mechanism is performed to update the video frames within the local window.

[0235] In some cases, self-attention mechanisms are applied to input feature sequences to enable the model to selectively emphasize context-relevant portions of the input data. The self-attention module calculates a similarity score between each token and all other tokens within the same input sequence, generating attention weights for reweighting and aggregating the input features. This mechanism allows for dynamic context modeling and improves feature representation by incorporating long-range relevance. Input sequences can include image features, video frames, or tokenized data, and the resulting attention-enhanced representations support tasks such as classification, generation, or alignment.

[0236] Attention mechanisms are a key component in some ANN architectures, enabling ANNs to selectively focus on different parts of the input sequence, thus assigning different levels of importance or attention to each part. Attention mechanisms achieve selective attention by considering the relevance of each input element to the current state of the ANN.

[0237] According to some aspects, an ANN employing an attention mechanism receives an input sequence and maintains a current state representing understanding or context. For each element in the input sequence, the attention mechanism calculates an attention score, which indicates the importance or relevance of that element given the current state. The attention score is transformed into attention weights through a normalization process (such as applying a softmax function). The attention weights represent the contribution of each input element to the overall attention. The attention weights are used to calculate a weighted sum of the input elements, resulting in a context vector. The context vector represents the information to which attention has been applied, or the portion of the input sequence that the ANN deems most relevant to the current step. The context vector is combined with the ANN's current state to provide additional information and influence the ANN's subsequent predictions or decisions.

[0238] By incorporating an attention mechanism, ANN dynamically allocates attention to different parts of the input sequence, allowing it to focus on relevant information and capture correlations across longer distances.

[0239] The description and accompanying drawings herein illustrate exemplary configurations and do not represent all embodiments within the scope of the claims. For example, operations and steps may be rearranged, combined, or otherwise modified. Furthermore, structures and devices may be represented in block diagram form to illustrate relationships between components and to avoid obscuring the described concepts. Similar components or features may have the same name but may have different reference numerals corresponding to different figures.

[0240] Some modifications to this disclosure may be apparent to those skilled in the art, and the principles defined herein can be applied to other variations without departing from the scope of this disclosure. Therefore, this disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope of the principles and novel features disclosed herein.

[0241] The described methods can be implemented or performed by devices including general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, a conventional processor, a controller, a microcontroller, or a state machine. A processor can also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration). Therefore, the functions described herein can be implemented in hardware or software and can be performed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, these functions can be stored on a computer-readable medium in the form of instructions or code.

[0242] Computer-readable media include both non-transitory computer storage media and communication media (including any media that facilitates the transfer of code or data). Non-transitory storage media can be any available medium that is accessible by a computer. For example, non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disk (CD) or other optical disc storage, magnetic disk storage, or any other non-transitory medium used to carry or store data or code.

[0243] Furthermore, the connecting components can be appropriately referred to as computer-readable media. For example, if code or data is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, or microwave signals), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology are included in the definition of media. Combinations of media are also included within the scope of computer-readable media.

[0244] In this disclosure and the following claims, the word “or” indicates an inclusive list, such that a list of, for example, X, Y, or Z means X or Y or Z or XY or XZ or YZ or XYZ. The phrase “based on” is also not used to indicate a closed set of conditions. For example, a step described as “based on condition A” could be based on both condition A and condition B. In other words, the phrase “based on” should be interpreted as meaning “at least partially based on.” Furthermore, the words “a” or “an” indicate “at least one.”

Claims

1. A method comprising: Obtain input prompts that describe the scenario; A video generation model is used to generate global tags based on the input cues by performing an attention process based on multiple frame tags; Using the video generation model, frame tags are generated by performing an attention process based on a subset of the global tags and the multiple frame tags in the local window; as well as Using the video generation model, a synthetic video is generated based on the frame markers, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

2. The method according to claim 1, further comprising: Text tags are generated based on the input prompts, wherein the global tags and the frame tags are generated by performing an attention process based on the text tags.

3. The method of claim 1, wherein generating the frame marker comprises: The attention process is performed based on multiple keyframes that correspond to multiple frame blocks, each of which includes a unique subset of the multiple frame tags.

4. The method of claim 1, wherein generating the frame marker comprises: Initialize multiple noisy frame markers; as well as The multiple noisy frame tags are denoised based on the global tag to obtain the multiple frame tags.

5. The method according to claim 1, further comprising: The global flag is updated iteratively based on the frame flag; as well as The frame marker is iteratively updated based on the global marker.

6. The method of claim 1, wherein generating the synthesized video comprises: Generate multiple image frames corresponding to the multiple frame tags.

7. The method according to claim 1, further comprising: Identify the diffusion time step; as well as Whether to perform selective attention on the frame marker is determined based on the diffusion time step, wherein the selective attention is based on the local window.

8. A non-transitory computer-readable medium storing code for video processing, the code including instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations including: Obtain input prompts that describe the scenario; Generate text tags based on the input prompts; Using a video generation model, frame tags are generated by performing an attention process based on the text tags and a subset of multiple frame tags in a local window; as well as Using the video generation model, a synthetic video is generated based on the frame markers, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

9. The non-transitory computer-readable medium of claim 8, wherein the code further comprises instructions, which, when executed by the at least one processor, cause the at least one processor to perform an operation, the operation comprising: Using the video generation model, a global marker is generated based on the input cue by performing an attention process based on the plurality of frame markers, wherein the frame markers are based on the global marker.

10. The non-transitory computer-readable medium of claim 8, wherein generating the frame marker comprises: The attention process is performed based on multiple keyframes that correspond to multiple frame blocks, each of which includes a unique subset of the multiple frame tags.

11. The non-transitory computer-readable medium of claim 9, wherein generating the frame marker comprises: Initialize multiple noisy frame markers; as well as The multiple noisy frame tags are denoised based on the global tag to obtain the multiple frame tags.

12. The non-transitory computer-readable medium of claim 9, wherein the code further comprises instructions, which, when executed by the at least one processor, cause the at least one processor to perform an operation, the operation comprising: The global flag is updated iteratively based on the frame flag; as well as The frame marker is iteratively updated based on the global marker.

13. The non-transitory computer-readable medium of claim 8, wherein generating the synthesized video comprises: Generate multiple image frames corresponding to the multiple frame tags.

14. The non-transitory computer-readable medium of claim 8, wherein the code further comprises instructions, which, when executed by the at least one processor, cause the at least one processor to perform an operation, the operation comprising: Identify the diffusion time step; as well as Whether to perform selective attention on the frame marker is determined based on the diffusion time step, wherein the selective attention is based on the local window.

15. A system comprising: Memory components; as well as A processing device coupled to the memory component, the processing device being configured to perform operations including: Obtain input prompts that describe the scenario; A video generation model is used to generate global tags based on the input cues by performing an attention process based on multiple frame tags; Using the video generation model, frame tags are generated by performing an attention process based on a subset of the global tags and the multiple frame tags in a local window; and Using the video generation model, a synthetic video is generated based on the frame markers, wherein the synthetic video depicts the scene and includes image frames corresponding to the frame markers.

16. The system according to claim 15, wherein: The video generation model includes a diffusion network.

17. The system of claim 15, further comprising: A text encoder configured to generate text markers based on the input prompt, wherein the global markers and the frame markers are generated by performing an attention process based on the text markers.

18. The system of claim 15, wherein generating the frame marker comprises: The attention process is performed based on multiple keyframes that correspond to multiple frame blocks, each of which includes a unique subset of the multiple frame tags.

19. The system of claim 15, wherein generating the frame marker comprises: Initialize multiple noisy frame markers; as well as The multiple noisy frame tags are denoised based on the global tag to obtain the multiple frame tags.

20. The system of claim 15, wherein generating the synthesized video comprises: Generate multiple image frames corresponding to the multiple frame tags.