Denoising noise markers using set of anchor markers to generate digital video

Generative AI digital vision system solves the problems of insufficient accuracy, efficiency and flexibility of existing systems in generating digital video by converting digital images into a set of image tags and using a diffusion converter model to process anchor tags and noise tags, thus achieving high-quality video generation.

CN121665082APending Publication Date: 2026-03-13ADOBE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510869199.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-23
Filing Date
2025-06-26
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing systems suffer from insufficient accuracy, efficiency, and operational flexibility when generating digital videos. They cannot accurately include the content of the input image in the video and consume excessive computing resources.

Method used

A generative AI digital vision system is used to generate high-quality digital video by converting digital images into a set of image tags and generating a set of anchor tags, and then using a diffusion converter model to process the anchor tags and noise tags.

Benefits of technology

It improves the accuracy and efficiency of video generation, reduces the consumption of computing resources, and enables flexible response to image-to-video requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121665082A_ABST
    Figure CN121665082A_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media for de-noising noise markers using a set of anchor markers to generate digital video based on the de-noised markers. In particular, the disclosed system generates a set of image tags from a digital image that is part of an image-to-video request. Further, the disclosed system generates a set of anchor markers from the set of image markers by adding a time step embedding to the set of image markers, the time step embedding indicating that the set of anchor markers is fully denoised. Further, the disclosed system generates a combined marker from a set of anchor markers and noise markers generated from noise. In addition, the disclosed system generates de-noised indicia by using a diffusion converter model to process the combined indicia. Further, according to the de-noising flag, the disclosed system generates a digital video including the digital image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the rights and priority of U.S. Provisional Application 63 / 693,660, filed September 11, 2024. The entire contents of the aforementioned application are incorporated herein by reference. Background Technology

[0003] Significant advancements have been seen in hardware and software platforms for performing generative tasks in recent years. In fact, systems offer a variety of approaches to generating still images and moving videos. For example, systems create different architectures for generating content in different modalities. Specifically, systems customize architectures for creating digital video from various input cues. Summary of the Invention

[0004] One or more embodiments described herein utilize systems, methods, and non-transitory computer-readable media that implement an artificial intelligence (AI) architecture for performing image-to-video requests to provide benefits and / or solve one or more problems in the art. For example, the disclosed system generates a set of image tokens for a digital image portion of an image-to-video request, and also generates a set of anchor tokens from the set of image tokens by adding a time step indicating that the set of image tokens has been fully denoised. Furthermore, in some embodiments, the disclosed system processes the anchor token set and the noise tokens using a diffusion transformer model to generate denoised tokens. Additionally, in some embodiments, the disclosed system generates digital video based on the denoised tokens, and the digital video includes at least a portion of the digital image as part of the image-to-video request.

[0005] Additional features and advantages of one or more embodiments of this disclosure are set forth in the detailed description below, and will be apparent in part from the detailed description, or may be learned by practice of such exemplary embodiments. Attached Figure Description

[0006] This disclosure will describe one or more embodiments of the invention with reference to the accompanying drawings, which include additional features and details. The following paragraphs briefly describe these drawings, in which:

[0007] Figure 1 The illustration shows an example environment in which a generative AI digital vision system based on one or more implementations operates;

[0008] Figure 2The illustration shows an example diagram of a generative AI digital vision system that receives an image-to-video request including a digital image and generates a digital video including the digital image, according to one or more implementations.

[0009] Figure 3A The illustration shows an example of a generative AI digital vision system that generates text tags from one or more conditional prompts, generates a set of anchor tags, and initializes noise tags based on one or more implementations.

[0010] Figure 3B The diagram illustrates an example of a generative AI digital vision system that generates text tags from one or more unconditional prompts, generates a set of anchor tags, and initializes noise tags based on one or more implementations.

[0011] Figure 4 The illustration shows an example of a generative AI digital vision system that generates a final output label by executing a first and second flow through a diffusion converter model according to one or more implementations.

[0012] Figure 5 The illustration shows an example of a generative AI digital vision system that uses a transformer block of a diffusion transformer model, based on one or more implementations, to generate conditionally labeled and unconditionally labeled outputs.

[0013] Figure 6 The illustration shows an example diagram of a generative AI digital vision system that converts the final labeled output into digital video according to one or more implementations;

[0014] Figure 7 The illustration shows an example of a generative AI digital vision system that receives training videos and generates image labels based on one or more implementations.

[0015] Figure 8 The illustration shows an example diagram of a generative AI digital vision system that adds spatial-temporal location encoding to image labeling according to one or more implementations;

[0016] Figure 9 The illustration shows an example of a generative AI digital vision system that modifies the parameters of a diffusion converter model based on one or more implementations of the generative loss metric.

[0017] Figure 10 The illustration shows a schematic diagram of a generative AI digital vision system based on one or more implementations;

[0018] Figure 11 The diagram illustrates a series of actions for generating digital video based on one or more implementations;

[0019] Figure 12An example of a diffusion model for denoising noisy data in a latent space, based on one or more implementations, is shown;

[0020] Figure 13 Examples of methods for media generation based on one or more implementations are shown;

[0021] Figure 14 The denoising process is illustrated according to one or more implementations;

[0022] Figure 15 A flowchart is shown illustrating a step-by-step process for training a machine learning model, based on one or more implementations.

[0023] Figure 16 A flowchart is shown for initializing an untrained model and updating its parameters according to one or more implementations;

[0024] Figure 17 Examples of computing devices based on one or more implementations are shown; and

[0025] Figure 18 An example of a generative AI digital vision system device (e.g., for interacting with generative AI digital vision system 102) based on one or more implementations is shown. Detailed Implementation

[0026] One or more embodiments described herein include frame anchoring techniques for generating high-quality and accurate digital video based on image-to-video requests. Existing systems performing generative tasks suffer from various problems related to accuracy, efficiency, and operational flexibility. Specifically, existing systems suffer from computational inaccuracies. For example, existing systems perform video generation; however, when performing generative tasks, they fail to accurately include content specified in the prompt within the generated digital video. For instance, existing systems utilize architectures that do not fully consider the information specified in the video generation request. Therefore, existing systems utilize generative architectures that generate inaccurate digital video.

[0027] Existing systems generate videos from user-provided prompts; however, they suffer from generating content lacking strong image / video and / or text semantic alignment (e.g., inaccurate digital videos). Furthermore, existing systems frequently generate low-quality and / or damaged frames that fail to capture the requested subject.

[0028] Furthermore, existing systems often struggle to accurately generate video from input images. For example, existing systems generate video based on input images but cannot include the input images as coherent frames within the video. Consequently, existing systems suffer from inaccurate video quality and low video quality when performing image-to-video tasks.

[0029] Related to accuracy issues, existing systems suffer from low computational efficiency. Specifically, existing systems typically require prompts and re-prompts to generate satisfactory digital video. In some instances, even with prompts and re-prompts, existing systems fail to generate accurate digital video. Thus, existing systems consume excessive computational resources and time to perform generative tasks.

[0030] Furthermore, conventional systems suffer from additional inefficiencies due to their complex converter-based architectures. Specifically, for conventional systems to generate video and image content, they typically require domain-specific complexity for the model architecture to capture all domain-specific data. Consequently, conventional systems demand significant time and resources to run models that generate content across domains. Moreover, existing systems suffer from operational inflexibility related to accuracy and computational efficiency. Specifically, existing systems exhibit rigid generative models that inaccurately and inefficiently generate digital video.

[0031] In some embodiments, generative AI digital vision systems overcome the shortcomings of existing systems. In one or more embodiments, the generative AI digital vision system implements frame anchoring techniques to generate digital video by converting a digital image (e.g., a portion of an image-to-video request) into a set of image tags and further into a set of anchor tags. For example, the generative AI digital vision system converts the set of image tags into a set of anchor tags by adding a time-step embedding that indicates the set of image tags has been fully denoised.

[0032] Accordingly, the generative AI digital vision system uses a diffusion converter model to process the anchor tag set (undenoised) and the noise tags to generate denoised tags. Specifically, the anchor tag set serves as a guide for removing noise from the noise tags, enabling the generative AI digital vision system 102 to generate digital video containing digital images as frames. In other words, the generative AI digital vision system utilizes the anchor tag set of the diffusion converter model to fully leverage the information presented in the digital images (e.g., conditions or anchor images) to remove noise from the noise tags.

[0033] In one or more embodiments, the generative AI digital vision system 102 uses a set of anchor tags to support the initial frame of the generated digital video. Specifically, the generative AI digital vision system 102 uses the set of anchor tags to ensure that the digital image ends at the initial frame of the generated digital video. In some embodiments, the generative AI digital vision system 102 uses the set of anchor tags to support the final frame, intermediate frames, any subset of frames, or any portion of frames within the generated digital video. Specifically, in some embodiments, the generative AI digital vision system 102 uses the set of anchor tags to ensure that the keyframes or motion frames of the generated digital video are digital images provided as part of an image-to-video request.

[0034] In one or more embodiments, the generative AI digital vision system 102 uses frame anchoring techniques to train a diffusion converter model. Specifically, the generative AI digital vision system receives training digital video comprising a sequence of frames. For example, the generative AI digital vision system converts the frame sequence into image tags and adds noise to the image tags of the frame sequence, in addition to frames used to adjust the diffusion converter model (e.g., anchor frames or adjustment images), to improve image-to-video generation during training and inference time. Accordingly, during inference time, the generative AI digital vision system demonstrates an improved generative capability to generate digital visual content using an artificial intelligence system.

[0035] In one or more embodiments, during the inference period, the generative AI digital vision system uses a trained diffusion converter model to perform high-quality and accurate image-to-video generation. Specifically, once the diffusion converter model is trained, the generative AI digital vision system receives an image-to-video request and performs multiple processes (e.g., a first process and a second process) on the diffusion converter model. For example, the generative AI digital vision system generates conditionally labeled outputs (e.g., outputs based on the conditional aspects of the image-to-video request) and unconditionally labeled outputs (e.g., outputs based on the unconditional aspects of the image-to-video request). Furthermore, the generative AI digital vision system combines the conditionally labeled outputs and the unconditionally labeled outputs to generate a final labeled output and further generates a digital video from the final labeled output.

[0036] As mentioned above, generative AI digital vision systems overcome the shortcomings of existing systems. For example, generative AI digital vision systems improve computational accuracy compared to existing systems. As mentioned above, existing systems suffer from the problem of inaccurately including content within the generated digital video. In contrast, generative AI digital vision systems generate a set of anchor tags (a set of image tags from digital images) to serve as a guide for removing noise from denoising tags. In doing so, generative AI digital vision systems more accurately include the content specified in the image-to-video request. For example, by keeping this set of image tags (for digital images, i.e., anchor images) fully denoised, generative AI digital vision systems use a diffusion-transformer model to fully consider the information presented in the digital images (e.g., conditions or anchor images) to remove noise from the noise tags.

[0037] Furthermore, compared to existing systems lacking strong image / video and / or text semantic alignment with image-to-video requests, in some embodiments, generative AI digital vision systems improve semantic alignment by identifying digital images (e.g., anchor images) to generate a set of anchor tags from them and further using the set of anchor tags to generate denoising tags. In other words, generative AI digital vision systems more accurately consider image-to-video cues to generate higher-quality digital video and higher-quality frames within the generated digital video.

[0038] Furthermore, compared to existing systems that typically cannot include input images (part of a visual cue) as coherent frames within the generated digital video, in some embodiments, the generative AI digital vision system receives an image-to-video request comprising digital images (e.g., digital images indicated as conditions for generating digital video) and generates a set of anchor tags from the digital images. Specifically, as mentioned, the generative AI digital vision system uses the set of anchor tags as guidance to remove noise from noise tags and create a digital video that includes the digital images as coherent frames.

[0039] Furthermore, in some embodiments, the generative AI digital vision system further improves the accuracy of existing systems by implementing unique training metrics. Specifically, the generative AI digital vision system uses training videos comprising frame sequences and adds noise to the image tags of the frame sequences in addition to the image tags of anchor frames (e.g., digital images). In doing so, the generative AI digital vision system optimizes the parameters of the diffusion converter model based on the set of anchor tags to learn to remove noise from the noise tags. Thus, the generative AI digital vision system improves the accuracy of existing systems by implementing frame anchoring techniques during training.

[0040] Furthermore, in some embodiments, the generative AI digital vision system improves accuracy by using spatial-temporal location encoding. Specifically, the generative AI digital vision system adds spatial-temporal location encoding to noise markers to further guide the diffusion converter model to remove noise from the noise markers.

[0041] In one or more embodiments, generative AI digital vision systems improve efficiency compared to existing systems. For example, existing systems require prompts and re-prompts to generate satisfactory digital video. In contrast, generative AI digital vision systems save computational resources by generating digital video that accurately conforms to image-to-video requests, thereby reducing the number of prompts and re-prompts from client devices.

[0042] Furthermore, in one or more embodiments, the generative AI digital vision system utilizes a single stream diffusion converter model, which simplifies complexity and improves the efficiency of generating digital video from digital images. Specifically, the generative AI digital vision system feeds input data, including a set of anchor tags, through a diffusion converter model (without additional modulation or adaptive normalization layers, hereinafter referred to as an adaLN layer), and the diffusion converter model treats the set of anchor tags as a guide to remove noise from noisy tags. Thus, in one or more embodiments, the generative AI digital vision system reduces the time and resources required to generate digital video from digital images.

[0043] As also mentioned above, in one or more embodiments, generative AI digital vision systems improve operational flexibility compared to existing systems. For example, generative AI digital vision systems provide dynamic flexibility to generate digital videos of varying ranges. For instance, generative AI digital vision systems allow for the precise inclusion of digital images in image-to-video requests for one or more frames in the generated digital video. Furthermore, generative AI digital vision systems improve flexibility by allowing image-to-video requests to specify whether visual cues (e.g., digital images) are included as initial frames, final frames, intermediate frames, portions of frames, one or more keyframes, and / or one or more motion frames.

[0044] Additional details about generative AI digital vision systems will now be provided with reference to the accompanying drawings. For example, Figure 1 The illustration shows a schematic diagram of an exemplary system environment 100 in which a generative AI digital vision system 102 operates. (See diagram for example.) Figure 1 As illustrated, system environment 100 includes (multiple) servers 104, a digital imaging system 106, a network 108, and client devices 110. Additionally, Figure 1 The illustration shows a digital imaging system 106 including a generative AI digital vision system 102, and the generative AI digital vision system 102 also includes a frame anchor system 105.

[0045] Although Figure 1 The system environment 100 is described as having a specific number of components, but the system environment 100 can have a different number of additional or alternative components (e.g., different numbers of servers, client devices, or other components communicating with the generative AI digital vision system 102 via network 108). Similarly, although Figure 1 The illustration shows a specific arrangement of (multiple) servers 104, network 108, and client devices 110, but various additional arrangements are possible.

[0046] (Multiple) servers 104, network 108, and client devices 110 directly or indirectly (e.g., through the following description) Figure 15 The network 108, discussed in more detail, is communicatively coupled to each other. Furthermore, the (multiple) servers 104 and client devices 110 include various computing devices (including, for example, those related to...). Figure 15 One or more computing devices (discussed in more detail).

[0047] As mentioned above, system environment 100 includes servers(s)104. In one or more embodiments, servers(s)104 process inputs for image-to-video requests or for training one or more artificial intelligence models or for generating video based on image-to-video requests. In one or more embodiments, servers(s)104 include data servers. In some implementations, servers(s)104 include communication servers or web hosting servers.

[0048] In one or more embodiments, client device 110 includes computing devices associated with one or more user accounts that submit image-to-video requests (e.g., media generation requests) to generative AI digital vision system 102 to generate media (e.g., based on text prompts and / or visual prompts). For example, generative AI digital vision system 102 trains one or more models from data using frame anchoring techniques (e.g., the diffusion converter model 107 portion of frame anchoring system 105).

[0049] In one or more embodiments, client device 110 includes a smartphone, tablet computer, desktop computer, laptop computer, head-mounted display device, or other electronic device. Client device 110 includes one or more software applications for generating content according to digital imaging system 106 (e.g., digital imaging application 112 includes a digital image editing application). In one or more embodiments, the digital imaging application includes a software application hosted on server(s) 104, which client device 110 can access via another application (such as a web browser).

[0050] To provide an exemplary implementation, in one or more embodiments, a generative AI digital vision system 102 on server(s)104 supports a generative AI digital vision system 102 on client device(s). For example, in some cases, a digital imaging system 106 on server(s)104 trains the generative AI digital vision system 102 (e.g., trains a generative model associated with frame anchor system 105, such as diffusion converter model 107) to be provided to client device(s)110 for implementation. In one or more embodiments, client device(s) obtain (e.g., download) the generative AI digital vision system 102 trained on server(s)104 for implementation. Once downloaded, the generative AI digital vision system 102 on client device(s)110 (e.g., trained on server(s)104) is provided with tools to instruct the generative AI digital vision system 102 to create media (e.g., generate digital video comprising frames provided by client device(s)110).

[0051] In an alternative implementation, the generative AI digital vision system 102 includes a web-hosted application that allows client device 110 to interact with content and services hosted on server(s) 104. In other words, client device 110 interacts with the generative AI digital vision system 102 without downloading it. For illustration, in one or more implementations, client device 110 accesses a software application supported by server(s) 104. In response, the generative AI digital vision system 102 on server(s) 104 provides tools for inputting instructions to generate digital video content (e.g., video with video captions and images).

[0052] Furthermore, in some implementations, the generative AI digital vision system 102 trains one or more AI models using a diffusion converter model to generate training embeddings and further utilizes these training embeddings to optimize the parameters of the diffusion converter model (e.g., the diffusion converter model implemented by the frame anchor system 105). Additionally, in one or more embodiments, the generative AI digital vision system 102 also generates improved positional encodings that capture spatial and temporal information of image blocks within frames of a frame sequence and use these improved positional encodings during inference and training periods (e.g., as data for removing guiding noise). For example, the generative AI digital vision system 102 utilizes positional encodings to further improve / optimize the parameters of the diffusion converter model.

[0053] In fact, in one or more embodiments, the generative AI digital vision system 102 is implemented wholly or partially by the various components of the system environment 100. For example, although Figure 1The illustration depicts a generative AI digital vision system 102 implemented or hosted on (multiple) servers 104, but different components of the generative AI digital vision system 102 can be implemented by various devices within the system environment 100. For example, one or more (or all) components of the generative AI digital vision system 102 may be implemented by computing devices or separate servers different from (multiple) servers 104. In fact, as... Figure 1 As shown, client device 110 includes generative AI digital vision system 102. The following will discuss... Figure 10 Example components to describe the generative AI digital vision system 102.

[0054] As mentioned above, in some embodiments, the generative AI digital vision system 102 generates digital video including digital images based on an image-to-video request. Figure 2 The illustration shows an overview diagram of a generative AI digital vision system 102 that utilizes a diffusion converter model to generate digital video from digital images, according to one or more embodiments. For example, Figure 2 An image-to-video request 200, including digital image 202, is shown.

[0055] In one or more embodiments, image-to-video request 200 refers to generative AI digital vision system 102 receiving a request to generate digital video. Specifically, generative AI digital vision system 102 receives image-to-video request 200 in the form of a cue from a client device to generate media conforming to that cue. For example, generative AI digital vision system 102 receives image-to-video request 200 as a visual cue (e.g., a digital image) and / or a text cue. For illustration, image-to-video request 200 includes specific parameters of generative AI digital vision system 102, such as creating digital video based on a provided digital image (e.g., a visual cue with digital image 202). Further, image-to-video request 200 may optionally include text cue to generate digital video, wherein the text cue specifies conditions (e.g., conditional cue) and unconditional cue (e.g., flexible settings included in the generated digital video), format, subject matter of the media, style of the media, mood or thematic content, and any additional details (e.g., aspect ratio, frames per second, lens size, camera angle, type of motion such as zoom in or zoom out, etc.).

[0056] As mentioned, the image-to-video request includes visual cues. In one or more embodiments, a visual cue refers to visual input that guides the generative AI digital vision system 102 to generate media. For example, a visual cue includes a digital image 202. Further, in some instances, the visual cue also includes a text cue along with the digital image 202. For illustration, the generative AI digital vision system 102 receives a visual cue that includes an image and a text cue describing the media to be generated.

[0057] In one or more embodiments, the digital image 202 includes various graphical elements. Specifically, the graphical elements include pixel values ​​that define the spatial and visual aspects of the digital image, such as text and image objects. For example, the digital image 202 is a rasterized image comprising a pixel grid. Specifically, the rasterized image includes a fixed resolution determined by a plurality of pixels within the digital image 202. Furthermore, in some embodiments, the digital image 202 is used as an anchor frame or adjustment frame for generating digital video.

[0058] like Figure 2 As shown, the generative AI digital vision system 102 uses a diffusion-to-video request 200 to process an image-to-video request 200 using a diffusion-to-transformer model 204 (e.g., including a transformer block 206 in some embodiments) to generate digital video 208. Specifically, the digital video 208 includes digital images 202 as frames within frames. In one or more embodiments, the generative AI digital vision system 102 generates digital video. As discussed in more detail below, in some embodiments, the generative AI digital vision system 102 utilizes digital video to train one or more models.

[0059] In one or more embodiments, digital video 208 refers to a media form that is encoded and stored in a digital format. Specifically, digital video 208 includes a sequence of frames (e.g., images, keyframes, and / or motion frames) and each frame in the sequence is displayed sequentially. For example, digital video 208 includes a specific resolution (480p, 720p, 1080p, 4K, 8K, etc.), which refers to the specific number of pixels displayed (e.g., the resolution of a video defines the sharpness and clarity of the digital video). Further, digital video 208 includes a frame rate (e.g., the number of frames displayed per second in the video, such as 24fps, 30fps, etc.), an aspect ratio (e.g., the width and height dimensions of the frames, such as 16:9 or 4:3), compression (e.g., the file size of the digital video), and audio that travels with the digital video (e.g., an audio file synchronized with the frames of the digital video).

[0060] In one or more embodiments, digital video 208 includes a sequence of frames. For example, a sequence of frames refers to a plurality of still images that are displayed sequentially to create motion-awareness. Specifically, each frame of the sequence represents a single moment in time, and when the sequence of frames is played together, it produces continuous motion and creates video content. In other words, the sequence of frames includes temporal continuity, where each frame in the sequence represents the next moment in time and motion is simulated as one moves from one frame to the next.

[0061] In one or more embodiments, digital video 208 includes image frames. For example, an image frame refers to a still image representing the content of digital video 208. Specifically, in one or more embodiments, generative AI digital vision system 102 considers the first frame of a frame sequence (e.g., frame zero) as an image frame. In other words, an image frame refers to the first visual element displayed at the beginning of the video (e.g., the beginning of a still image in the video).

[0062] In one or more embodiments, a keyframe refers to an image frame that stores visual data of the start or end of an action or the position of an object or character. Specifically, the video includes multiple keyframes. In other words, the generative AI digital vision system 102 utilizes keyframes as complete image frames that act as anchor points for motion. For illustration, the video includes a sequence of frames, and this sequence includes keyframes every 16 frames.

[0063] In one or more embodiments, digital video 208 includes at least one motion frame. For example, generative AI digital vision system 102 utilizes motion frames as intermediate frames between keyframes to store changes or differences from previous frames. Specifically, generative AI digital vision system 102 utilizes motion frames to store information related to changes between consecutive frames, such as changes in the position or color of an object from one frame to the next. Further, generative AI digital vision system 102 utilizes motion frames concatenated with keyframes during playback of digital video 208 to create a perception of smooth motion from one keyframe to the next.

[0064] As mentioned above, the generative AI digital vision system 102 generates a set of anchor tags. Figures 3A to 3B A generative AI digital vision system 102 is illustrated, which generates a set of anchor tags and text tags for conditional and unconditional prompts according to one or more embodiments. For example, Figures 3A to 3B A generative AI digital vision system 102 is shown generating an image embedding 306 from a digital image 300 using an image encoder 302 (e.g., an encoder of a dual VAE model mentioned below). Specifically, the digital image 300 acts as an anchor digital image / modulated digital image as part of an image-to-video request.

[0065] In one or more embodiments, the image encoder 302 is a neural network (or one or more layers of a neural network) that extracts features associated with a digital image. In some cases, the image encoder 302 refers to a neural network that extracts features from the digital image 300 and encodes the features from the digital image 300. For example, the image encoder 302 includes a specific number of layers that include one or more fully connected and / or partially connected layers of neurons that extract image patches from the digital image 300 and encode localized features of the digital image 300. For illustration, in one or more embodiments, the generative AI digital vision system 102 generates an image embedding 306 representing a complete frame of a digital image.

[0066] In one or more embodiments, the generative AI digital vision system 102 utilizes an image encoder 302 to generate embeddings (e.g., image embeddings 306). In some embodiments, the embedding includes a numerical representation (e.g., a vector) of the digital image 300. For example, the embedding captures features and characteristics of the digital image 300. For illustration, the embedding includes semantic information such as the presence, shape, and spatial relationships of objects.

[0067] also, Figures 3A to 3B The illustration shows a generative AI digital vision system 102 that uses a tokenization model 308 to generate an image tag set 310. In one or more embodiments, the generative AI digital vision system 102 converts embeddings (e.g., image embeddings) into image tags (e.g., visual tags).

[0068] For example, the generative AI digital vision system 102 utilizes a tokenization model 308 to patch the image embedding 306. Specifically, the tokenization model 308 transforms the embedding into smaller tiles or grids, which are treated as individual labels for further processing (e.g., adding noise and then denoising). For example, the generative AI digital vision system 102 utilizes tiles to efficiently process high-dimensional image data. To illustrate, the generative AI digital vision system 102 flattens each tile of the embedding (e.g., flattens it into a one-dimensional vector), transforms the flattened tiles into lower-dimensional representations, and maps the flattened lower-dimensional tiles to fixed-length feature vectors.

[0069] Accordingly, the generative AI digital vision system 102 processes flattened, fixed-length feature vectors as image tags and utilizes a diffusion converter model to process these image tags. Furthermore, in some embodiments, the generative AI digital vision system 102 adds positional encoding to each block (e.g., image tag) to encode spatial information about where that block belongs in the digital image.

[0070] In one or more embodiments, the generative AI digital vision system 102 selects a set of image blocks from the digital image 300. Specifically, the generative AI digital vision system 102 generates the set of image blocks by subdividing the digital image into smaller regions. For example, the generative AI digital vision system 102 subdivides the digital image 300 into blocks based on a predetermined resolution (e.g., 256×256), where each block represents a local region within the digital image 300. In some embodiments, image blocks in the set of image blocks do not share any pixel values ​​with other image blocks. In some embodiments, the pixel values ​​of image blocks in the set of image blocks are superimposed with those of neighboring image blocks. Accordingly, in one or more embodiments, the generative AI digital vision system 102 subdivides the digital image 300 into image blocks, wherein some image blocks in the image blocks do not have their pixel values ​​superimposed with those of other image blocks, and some image blocks in the image blocks have their pixel values ​​superimposed with those of other image blocks.

[0071] As mentioned above, the generative AI digital vision system 102 utilizes a tokenization model 308 to generate image tags. In one or more embodiments, this image tag set 310 refers to embedded image tags from digital images (e.g., digital images of visual cues). For example, the image tags in the image tag set 310 come from image blocks of the digital image 300. In other words, the generative AI digital vision system 102 uses the tokenization model 308 to decompose the digital image 300 into image blocks and further converts each image block into an image tag. Specifically, the generative AI digital vision system 102 generates the image tag set 310 to be used as anchor tags in the denoising process.

[0072] like Figures 3A to 3BAs shown, the generative AI digital vision system 102 adds a time-step embedding 313 to the image tag set 310. In one or more embodiments, the generative AI digital vision system 102 adds the time-step embedding 313 to the tags (e.g., the image tag set 310). For example, the time-step embedding 313 refers to an embedding that represents a specific amount of noise added to the tags / tag set at a particular time step. In other words, the generative AI digital vision system 102 generates a time-step embedding 313 corresponding to a first converter block, a second time-step embedding corresponding to a second converter block, and a third time-step embedding corresponding to a third converter block. For example, the time-step embedding 313 indicates a particular time step in which noise is added to noise tags, such that the generative AI digital vision system 102 determines how much noise is removed from the tags at a particular converter block. Furthermore, in some embodiments, the generative AI digital vision system 102 adds the time-step embedding 313 to the tag set, wherein the time-step embedding indicates that the tag set is completely denoised. In other words, in some instances, the time step embedding instructs the generative AI digital vision system 102 not to perform any denoising on a particular set of image labels.

[0073] like Figures 3A to 3B As further illustrated, the generative AI digital vision system 102 adds a time-step embedding 313 to the image tag set 310 to generate an anchor tag set 315. In one or more embodiments, the anchor tag set 315 refers to the time-step embedding 313 added to the image tag set 310. Specifically, the anchor tag set 315 refers to the conditioning input of the diffusion converter model.

[0074] In other words, the anchor tag set 315 refers to the anchor / guide used by the generative AI digital vision system to guide the denoising / removal of noise from the noise tags. As mentioned above, the image tag set 310 corresponds to the digital image 300, which is part of the visual cue. Therefore, during inference, the generative AI digital vision system 102 uses the anchor tag set 315 to ensure that the output (e.g., the generated digital video) includes the digital image 300 from the visual cue. Specifically, the generative AI digital vision system 102 denoises the noise tags according to the anchor tag set 315. In other words, the generative AI digital vision system 102 anchors its generative mechanism to create a digital video that includes fully denoised content (e.g., the anchor tag set 315).

[0075] like Figure 3AAs shown by the dashed lines, in some embodiments, the generative AI digital vision system 102 receives a conditional cue 304 from a client device. For example, the generative AI digital vision system 102 implicitly or explicitly requests the receipt of the conditional cue 304 from an image to a video. Specifically, the generative AI digital vision system 102 implicitly receives the conditional cue 304 by receiving a digital image 300, and the generative AI digital vision system 102 assumes that the digital image 300 is a condition for generating the digital video. In some embodiments, the generative AI digital vision system 102 explicitly receives the conditional cue 304 as part of a text prompt.

[0076] In one or more embodiments, the generative AI digital vision system 102 receives an image-to-video request as a text prompt. Specifically, the generative AI digital vision system 102 receives a text prompt from a client device, which describes in text form the content to be included in the digital video generated by the generative AI digital vision system 102. For example, the text prompt describes specific parameters to be included in the media generated by the generative AI digital vision system 102.

[0077] In one or more embodiments, condition cue 304 refers to the generative AI digital vision system 102 generating data based on specific input conditions. Specifically, condition cue 304 refers to a specific instruction guiding the generation process to produce an output aligned with a given context. For example, in some embodiments, a digital image serves as condition cue 304. In other words, the generative AI digital vision system 102 receives a digital image as a condition for generating digital video (e.g., the digital video must include digital images as one or more frames).

[0078] In one or more embodiments, the generative AI digital vision system 102 receives conditional prompts 304 as part of an instruction from a client device (e.g., a checkbox indicating that the uploaded digital image must be part of the generated digital video). In some embodiments, the generative AI digital vision system 102 receives conditional prompts as part of an instruction that includes both textual and visual prompts. For example, the generative AI digital vision system 102 receives a digital image and further receives the text instruction “Generate a digital video of a car driving towards the city at night” or “Generate a digital video of a car driving towards the city at night, like the image just provided.”

[0079] like Figure 3B As shown by the dashed lines, in some embodiments, the generative AI digital vision system 102 also receives an unconditional cue 316. Similar to the conditional cue 304, in some embodiments, the generative AI digital vision system 102 receives the unconditional cue 316 as an explicit or implicit instruction.

[0080] In one or more embodiments, unconditional cue 316 refers to the generative AI digital vision system 102 generating data without conditioning or guidance. Specifically, the generative AI digital vision system 102 does not rely on specific inputs or cues to guide output generation (e.g., digital video). In other words, unconditional cue 316 does not constrain the generation process utilizing external data (e.g., digital images). For illustration, the generative AI digital vision system 102 receives unconditional cue 316 as part of an image-to-video request (e.g., generating a video with lower resolution and poorer aesthetics).

[0081] like Figures 3A to 3B As shown, the generative AI digital vision system 102 also utilizes a text encoder 312 to process conditional cues 304 and unconditional cues 316. In one or more embodiments, the generative AI digital vision system 102 utilizes a text encoder 312 to process text cues. Specifically, the text encoder includes components of a neural network that convert text data (e.g., text cues) into numerical representations.

[0082] For example, the generative AI digital vision system 102 utilizes a text encoder 312 to convert text prompts into text encodings (e.g., text tags). Further, the generative AI digital vision system 102 utilizes the text encoder in various ways. For example, the generative AI digital vision system 102 utilizes the text encoder to i) determine the frequency of individual words in the text prompt (e.g., each word becomes a feature vector); ii) determine the weight of each word within the text prompt to generate text vectors that capture the importance of words within the text prompt; iii) generate low-dimensional text vectors in a continuous vector space representing words within the text prompt; and / or iv) generate contextualized text vectors by determining semantic relationships between words within the text prompt.

[0083] In one or more embodiments, the generative AI digital vision system 102 generates text tags 314 from conditional cues 304 and text tags 320 from unconditional cues 316. For example, the generative AI digital vision system 102 utilizes a text encoder 312 to generate representations of text cues for machine learning tasks. Specifically, individual text tags refer to words, subwords, or characters (e.g., "the", "on", "cat", "t", "showcasing", "show", "casing", etc.). Furthermore, the generative AI digital vision system 102 generates tags that represent specific meanings or purposes, such as the beginning or end of a sentence.

[0084] As mentioned above, the generative AI digital vision system 102 utilizes multiple processes through a diffusion converter model to generate the final output label. Figure 4The illustration depicts a generative AI digital vision system 102, according to one or more embodiments, performing a first process and a second process to generate a final output label. For example, Figure 4 A generative AI digital vision system 102 is shown that utilizes a diffusion converter model 410 to generate a final labeled output 418.

[0085] In one or more embodiments, a machine learning model includes a computer algorithm or a collection of computer algorithms trained and / or tuned based on inputs to approximate an unknown function. For example, a machine learning model includes a computer algorithm with branches, weights, or parameters that are modified based on training data to improve for a specific task. Thus, machine learning models utilize one or more learning techniques to improve accuracy and / or effectiveness. Example machine learning models include various types of decision trees, support vector machines, Bayesian networks, random forest models, or neural networks (e.g., deep neural networks).

[0086] Similarly, neural networks include machine learning models of interconnected artificial neurons (e.g., organized into layers) that communicate and learn approximate complex functions and generate outputs based on multiple inputs provided to the model. In some instances, neural networks include algorithms (or sets of algorithms) that implement deep learning techniques that utilize a set of algorithms to model high-level abstractions in data. For illustration, in some embodiments, neural networks include convolutional neural networks, recurrent neural networks (e.g., long short-term memory neural networks), transducer neural networks, generative inverse neural networks, graph neural networks, diffuse neural networks, or multilayer perceptrons. In some embodiments, neural networks include neural networks or combinations of neural network components.

[0087] In one or more embodiments, the generative AI digital vision system 102 utilizes a diffusion model as a neural network. For example, a diffusion model refers to a generative machine learning model that reconstructs data by removing noise from input data. Specifically, the generative AI digital vision system 102 trains the diffusion model to remove noise, compares the denoised representation with the ground truth, and modifies the parameters of the diffusion model.

[0088] In one or more embodiments, the generative AI digital vision system 102 utilizes a diffusion converter model. Specifically, a diffusion converter model refers to a model architecture that utilizes the principles of a diffusion model based on a converter architecture. For example, a diffusion converter model includes a deep learning self-attention mechanism for processing sequential data. For instance, a diffusion converter model uses a self-attention mechanism to establish relationships between elements in a sequence. To illustrate, the generative AI digital vision system 102 utilizes a diffusion converter model to denoise noisy representations (e.g., noise labels) to reconstruct data and generate media (e.g., videos, images, text, etc.).

[0089] In one or more embodiments, the first process refers to the generative AI digital vision system 102 passing initial input through a diffusion converter model 410. Specifically, the generative AI digital vision system 102 generates a conditional label output 412 from the first process.

[0090] Figure 4 The illustration shows a generative AI digital vision system 102 that generates combined tags 401. For example, Figure 4 A combined marker 401 is shown, which includes text markers 400 from conditional prompts, a set of anchor markers 402, and noise markers 404. Specifically, during the inference period, the generative AI digital vision system 102 initializes the noise markers 404.

[0091] In one or more embodiments, the generative AI digital vision system 102 initializes noise markers 404 from a noise distribution to generate noise markers 404. Specifically, during inference epochs (e.g., runtime), the generative AI digital vision system 102 utilizes a diffusion converter model 410 to process the noise markers 404. For example, the generative AI digital vision system 102 initializes the noise markers 404 and processes the noise markers 404 along with an anchor marker set 402. For example, the generative AI digital vision system 102 generates noise markers 404 by adding Gaussian noise sampled from a normal distribution having zero mean and a specified standard deviation, wherein the noise distribution ranges from t=0 to t=1000, t=1000 indicating that the markers are completely noised, and t=0 indicating that the markers are completely denoised.

[0092] In one or more embodiments, combined marker 401 refers to the generative AI digital vision system 102 combining (e.g., cascading) anchor marker set 402 and noise marker 404 to pass through a diffusion converter model. Specifically, combined marker 401 refers to the generative AI digital vision system 102 generating combined input for a first process through the diffusion converter model 410. In some embodiments, the generative AI digital vision system 102 also combines text marker 400 with anchor marker set 402 and noise marker 404 to generate combined marker 401. Specifically, the generative AI digital vision system 102 generates combined marker 401 from text markers of conditional prompts for image-to-video requests.

[0093] like Figure 4As illustrated, the generative AI digital vision system 102 generates a conditional tag output 412 from a first process via a diffusion converter model 410. In one or more embodiments, the conditional tag output 412 refers to a denoised tag generated by the generative AI digital vision system 102 that processes the combined tag 401. Specifically, the generative AI digital vision system 102 uses the combined tag 401 to generate a denoised tag based on the anchor tag set 402 (e.g., the denoising process is guided by a digital image). Therefore, the conditional tag output 412 refers to data indicating the conditional aspects of an image-to-video request.

[0094] Furthermore, as shown, the generative AI digital vision system 102 also performs a second process to generate an unconditional labeled output 414. In one or more embodiments, the second process refers to the generative AI digital vision system 102 passing a second input through a diffusion converter model 410. Specifically, the generative AI digital vision system 102 generates the unconditional labeled output 414 from the second process. For example, the generative AI digital vision system 102 generates the unconditional labeled output 414 from additional combined labels 403.

[0095] In one or more embodiments, the additional combined marker 403 refers to the generative AI digital vision system 102 combining (e.g., cascading) the anchor marker set 402 and the noise marker 408 for use in a second process through the diffusion converter model 410. In some embodiments, the generative AI digital vision system 102 generates the additional combined marker 403 by combining the text marker 406 of the unconditional cue from the image-to-video request, the anchor marker set 402, and the noise marker 408 (e.g., a noise marker initialized for the combined marker 401 or an additional noise marker containing the same noise level as or different from the noise marker 404).

[0096] Although the additional combined marker 403 is used in the second process and the text marker 406 from the unconditional prompt is used, in one or more embodiments, the generative AI digital vision system 102 uses the anchor marker set 402 in the second process. By using the anchor marker set 402 in both the first and second processes, the generative AI digital vision system 102 improves the accuracy and quality of the generated digital video (e.g., particularly if the goal is to include digital images within the generated digital video).

[0097] like Figure 4As shown, the generative AI digital vision system 102 also utilizes a diffusion converter model 410 to generate an unconditional labeled output 414 from additional combined labels 403. Specifically, the generative AI digital vision system 102 generates the unconditional labeled output 414 based on unconditional cues and is further guided by the anchor label set 402 during the denoising process. Therefore, the unconditional labeled output 414 indicates the unconditional aspect of the image-to-video request.

[0098] Figure 4 The generative AI digital vision system 102 is also illustrated to generate a final labeled output 418 from conditionally labeled output 412 and unconditionally labeled output 414. In one or more embodiments, the final labeled output 418 refers to a combination of the conditionally labeled output 412 and the unconditionally labeled output 414. Specifically, the final labeled output 418 refers to the generative AI digital vision system 102 using a classifier-free guided model 416 to facilitate a diffusion converter model 410 in generating digital video based on either the conditionally labeled output 412 or the unconditionally labeled output 414. For example, the generative AI digital vision system 102 uses the final labeled output 418 to generate digital video.

[0099] In one or more embodiments, the classifier-free guided model 416 refers to a model that does not rely on a fast classifier. Specifically, the generative AI digital vision system 102 uses the classifier-free guided model 416 to steer the output of the diffusion converter model 410 toward desired characteristics. For example, the generative AI digital vision system 102 uses weights to indicate whether to support the conditional output or the unconditional output more. In some embodiments, the generative AI digital vision system 102 uses the classifier-free guided model 416 to interpolate between the conditionally labeled output 412 and the unconditionally labeled output 414. Specifically, this interpolation includes guided scaling that facilitates the diffusion converter model to generate digital video based on the conditionally labeled output.

[0100] As discussed above, the generative AI digital vision system 102 utilizes an anchor mark set 402 to guide noise removal and anchor a digital image to a frame within the generated digital video. In some embodiments, the anchor mark set 402 corresponds to the initial frame of the digital video. For example, an image-to-video request includes a digital image and also indicates that the generated product should include the digital image as the initial frame of the digital video.

[0101] In some embodiments, the anchor tag set 402 corresponds to intermediate frames of a digital video. For example, an image-to-video request includes a digital image and also indicates that the generated product should include the digital image as one or more intermediate frames in the video. In some embodiments, the anchor tag set 402 corresponds to the final frame of a digital video. For example, an image-to-video request includes a digital image and also indicates that the generated product should include the digital image as the last or final frame of the video.

[0102] In some embodiments, the anchor tag set 402 corresponds to keyframes of a digital video. For example, an image-to-video request includes a digital image and also indicates that the generated product should include the digital image as one or more keyframes. In some embodiments, the anchor tag set 402 corresponds to motion frames of a digital video. For example, an image-to-video request includes a digital image and also indicates that the generated product should include the digital image as one or more motion frames.

[0103] In some embodiments, the anchor tag set 402 corresponds to a portion of a frame in a digital video. For example, an image-to-video request includes a digital image and also indicates that the generated product should include the digital image as part of a frame (e.g., a first, middle, last, keyframe, or motion frame). In other words, the generative AI digital vision system 102 uses digital images to cover a portion of a frame in the generated digital video and uses additional content to fill in the remainder of that frame.

[0104] As mentioned above, the generative AI digital vision system 102 utilizes a diffusion converter model with a streamlined architecture. Figure 5 The illustration depicts a generative AI digital vision system 102 that utilizes a converter block of a diffusion converter model to generate conditionally labeled outputs and unconditionally labeled outputs according to one or more embodiments. For example, Figure 5 A generative AI digital vision system 102 is shown processing a combination of tags 501, which includes text tags 500 (e.g., conditional cue from generating conditional tag output or unconditional cue from generating unconditional tag output), a set of anchor tags 502, and noise tags 504.

[0105] like Figure 5As shown, the generative AI digital vision system 102 utilizes a first converter block 506 of a diffusion converter model to process the combined marker 501. In one or more embodiments, a converter block refers to a single block within a single stream converter. Specifically, the generative AI digital vision system 102 utilizes a converter block of a single stream converter to remove noise from the noisy marker. For example, for a single stream converter with multiple converter blocks, the generative AI digital vision system 102 utilizes the first converter block to remove some noise from the noisy marker to generate an intermediate denoised marker.

[0106] In one or more embodiments, the generative AI digital vision system 102 utilizes a converter block to remove at least some noise from the noise markers 504. For example, the generative AI digital vision system 102 generates intermediate denoised markers using a first converter block 506. Specifically, intermediate denoised markers refer to partial noise markers. Specifically, once the generative AI digital vision system 102 has removed all noise from the noise markers using a single stream converter, the generative AI digital vision system 102 generates denoised markers (e.g., conditional / unconditional marker output 524).

[0107] Figure 5 The first converter block 506 is shown to include a self-attention layer 508, a combined self-attention layer output 510, a multilayer perceptron 512, and a multilayer perceptron output 514. In one or more embodiments, the self-attention layer 508 refers to a layer that captures the importance of different tags (e.g., words or chunks) in a sequence relative to each other. Specifically, the generative AI digital vision system 102 utilizes the self-attention layer 508 to capture relationships and dependencies between tags (e.g., for short-range and long-range dependencies).

[0108] In other words, the generative AI digital vision system 102 uses a self-attention layer 508 to determine how much attention a tag should give to another tag. To illustrate, the generative AI digital vision system 102 uses the self-attention layer 508 to generate three vectors for each tag: 1) a query vector (e.g., representing a tag that seeks information from other tags); 2) a key vector (e.g., representing a tag that provides information to other tags); and 3) a value vector (e.g., representing the actual content of the tag).

[0109] In one or more embodiments, the generative AI digital vision system 102 utilizes a self-attention layer 508 to generate a self-attention layer output that represents an updated set of intermediate noise tags (e.g., or denoised tags) incorporating information from other noise tags (e.g., the updated set of noise tags represents relationships between tags). In one or more embodiments, the generative AI digital vision system 102 also combines the self-attention layer output with an initial input to a converter block corresponding to the self-attention layer to generate a combined self-attention layer output 510.

[0110] In one or more embodiments, the generative AI digital vision system 102 utilizes a multilayer perceptron 512. For example, a multilayer perceptron 512 refers to an artificial neural network having a fully connected multilayer of neurons. Specifically, the multilayer perceptron 512 includes an input layer (where input data is fed into the network), hidden layers (e.g., an intermediate layer between the input and output layers, where the hidden layers receive input from all neurons in the previous layers), and an output layer that generates the multilayer perceptron output 514 (e.g., by combining the output of the self-attention layer with the output from the multilayer perceptron 512).

[0111] Figure 5 The generative AI digital vision system 102 is shown generating a first set 516 of intermediate denoised tags and utilizing a second converter block 518 to further generate a second set 520 of intermediate denoised tags. Furthermore, Figure 5 A generative AI digital vision system 102 is shown that utilizes the Nth converter block 522 to generate conditional / unconditional tag output 524.

[0112] Figure 6 A generative AI digital vision system 102 is illustrated, which converts a final tagged output into digital video according to one or more embodiments. For example, Figure 6 A generative AI digital vision system 102 is shown that uses a de-labeling model 602 to process the final labeled output 600.

[0113] In one or more embodiments, the generative AI digital vision system 102 transforms the final labeled output 600 (e.g., denoised labels) into embeddings by utilizing a de-labeling model 602. For example, the generative AI digital vision system 102 utilizes the de-labeling model 602 to integrate the denoised labels. Specifically, de-blocking involves the reverse process of reconstructing an image (e.g., a frame sequence) from a set of denoised labels (e.g., the final labeled output 600).

[0114] For example, the generative AI digital vision system 102 rearranges the denoised tags (e.g., final tag output 600) and combines the rearranged denoised tags into an initial (original) image structure / frame. In other words, the generative AI digital vision system 102 uses the de-tag model 602 to rearrange the tags to resemble the embedding 604 (e.g., putting the entire frame together).

[0115] Furthermore, in some embodiments, the generative AI digital vision system 102 utilizes a decoder 606 to process denoised tags (which have been de-blocked) and generate media items such as digital video 608. In one or more embodiments, the generative AI digital vision system 102 utilizes a decoder 606 comprising one or more layers (e.g., linear transformation, self-attention layers, softmax layers, etc.) to transform the embedding 604 into digital video 608. Specifically, the decoder 606 transforms the denoised tags in the latent space into images / frames in the pixel space. In one or more embodiments, the generative AI digital vision system 102 utilizes one or more decoders of a bivariate autoencoder model described in application 18 / 930,665, filed October 29, 2024, entitled “DUAL-VAE FOR MORE EFFICIENT AND EFFECTIVE DIFFUSION MODELTRAINING,” the entire contents of which are incorporated herein by reference.

[0116] As mentioned above, the generative AI digital vision system 102 trains the diffusion converter model in a way that more accurately includes digital images in the generated digital video. Figure 7 A generative AI digital vision system 102 is illustrated, which receives training digital video and generates image labels (e.g., a training label set) from the training digital video according to one or more embodiments. For example, Figure 7 The diagram illustrates that during the training phase, the generative AI digital vision system 102 receives a training digital video comprising a sequence of frames (e.g., a training frame sequence). For illustration, the training frame sequence of the training digital video comprises one hundred image tags, wherein the first ten tags correspond to the first frame, and the next ninety image tags correspond to the next nine frames of the frame sequence (e.g., ten tags per frame).

[0117] Figure 7 The illustration shows a generative AI digital vision system 102 that uses encoder 708 to generate embeddings 710 for the first frame 702, 720 for the second frame 704, and 730 for the Nth frame 706. Furthermore, Figure 7A generative AI digital vision system 102 is shown that uses a tokenization model 712 to generate image tags 714 for the first frame 702, image tags 724 for the second frame 704, and image tags 732 for the Nth frame 706.

[0118] In one or more embodiments, during training, the generative AI digital vision system 102 adds noise to image markers (e.g., clean markers corresponding to frames of a training video) at several time steps. For example, the generative AI digital vision system 102 adds noise to the markers at multiple time steps corresponding to multiple converter blocks (e.g., denoising blocks) in a diffusion converter model. Specifically, the generative AI digital vision system 102 randomly samples from a noise distribution to determine how much noise (t ranging from 0 to 1000) to add to the clean image markers.

[0119] In some embodiments, the generative AI digital vision system 102 adds the same amount of noise to all image labels during training, while in other embodiments, the generative AI digital vision system 102 varies the amount of noise added to the labels. Specifically, Figure 7 The generative AI digital vision system 102 is shown adding noise to image labels 724 to generate noise labels 726 and adding noise to image labels 732 to generate noise labels 734 (e.g., noise training labels). However, as... Figure 7 As illustrated, the generative AI digital vision system 102 does not add any noise to the image marker 714 (e.g., as indicated by t=0) because the image marker 714 comes from the first frame 702 that is being used as a condition / anchor frame (e.g., training anchor marker).

[0120] As further illustrated, the generative AI digital vision system 102 adds time step embedding 716 to image marker 714, adds time step embedding 728 to image marker 724, and adds time step embedding 736 to image marker 732. As discussed above, the time step embedding indicates to the generative AI digital vision system 102 utilizing the diffusion converter the amount of noise added to the image markers. Therefore, as Figure 7 As shown, t=T indicates that noise markers 726 and 734 are completely noise-enhanced.

[0121] As discussed above in the context of the inference period, the generative AI digital vision system 102 utilizes a set of anchor tags (e.g., image tags 714 without added noise, used as training anchor tags) to remove noise from noise tags. In one or more embodiments, during the training period, the generative AI digital vision system 102 also uses a set of anchor tags (e.g., image tags 714 without added noise, as indicated by the time step embedding 716) to remove noise from noise tags 726 and 734. Specifically, the noise tags are derived from a sequence of frames (e.g., excluding anchor frames, which are the frames used to create the set of anchor tags during the training period). Accordingly, the generative AI digital vision system 102 uses the set of anchor tags (e.g., image tags 714) to guide the denoising process during the training period.

[0122] also, Figure 7 It is shown that in one or more embodiments, the generative AI digital vision system 102 also utilizes text prompts 738 during training. Specifically, Figure 7 The text prompt 738 shows "a car driving towards the city at night". Additionally, Figure 7 A generative AI digital vision system 102 is shown that utilizes a text encoder 740 to generate text tags 742.

[0123] In one or more embodiments, the generative AI digital vision system 102 also staggers anchor frames across multiple frames in a frame sequence. Specifically, for a frame sequence comprising 50 frames, the generative AI digital vision system 102 utilizes the first, eleventh, twenty-first, thirty-first, and forty-first frames as anchor frames for training purposes. For example, the generative AI digital vision system 102 optimizes a diffusion converter model to process a subset of frames as anchor frames for removing noise from noise markers.

[0124] Figure 8 Further illustration is provided of a generative AI digital vision system 102 that incorporates spatial-temporal location encoding into a set of noise markers and / or anchor markers, according to one or more embodiments. For example, Figure 8 A generative AI digital vision system 102 is shown that generates a frame N embedding 804 for frame N 800 using an encoder 802. Furthermore, Figure 8 A generative AI digital vision system 102 is illustrated, which generates image tags 808 from frame N embeddings 804 using a tokenization model 806. Furthermore, in some embodiments, the generative AI digital vision system 102 generates temporal embeddings 810 and spatial embeddings 812 for image tags 814 of the image tags 808.

[0125] In one or more embodiments, spatial embedding 812 refers to a representation of the spatial relationships and positions of visual elements within frames (e.g., images) of a frame sequence. Specifically, spatial embedding 812 includes indications of where objects / elements are located in a frame, their orientation, size, and spatial relationships with different regions of the frame in which they are located. For example, generative AI digital vision system 102 utilizes coordinate information of objects / elements within a frame (e.g., x-dimension and y-dimension, and in some embodiments, z-dimension).

[0126] In some embodiments, spatial embedding 812 indicates an absolute position within a frame, and in some embodiments, spatial embedding 812 indicates a relative position (e.g., relative to other objects / elements within the frame). In one or more embodiments, the generative AI digital vision system 102 utilizes a centered two-dimensional coordinate graph to generate the spatial embedding 812 of the image label 814 when noise 816 is added to the image label 814.

[0127] In one or more embodiments, temporal embedding 810 refers to a representation of frames within a visual frame sequence. Specifically, generative AI digital vision system 102 utilizes temporal embedding 810 to capture motion information, action sequences, and transitions between frames within a frame sequence. In other words, generative AI digital vision system 102 generates temporal embedding 810 to create a representation of the sequential dependencies between frames in a frame sequence.

[0128] In one or more embodiments, the generative AI digital vision system 102 generates a temporal embedding 810 based on a timestamp and an inverse timestamp. For example, the generative AI digital vision system 102 determines the timestamp of image marker 814 (e.g., the first frame of a video frame sequence). Specifically, the timestamp of the first frame refers to a specific point in time within the entire video or frame sequence where the frame of image marker 814 appears relative to the beginning of the video.

[0129] Furthermore, the generative AI digital vision system 102 determines an inverse timestamp, which refers to the difference between the total length of the video and the time position (e.g., current position) of the frame of image tag 814 relative to the frame sequence. Additionally, the generative AI digital vision system 102 combines the timestamp and inverse timestamp of image tag 814 with noise 816 to generate a temporal embedding 810.

[0130] For illustration, the generative AI digital vision system 102 utilizes the methods discussed in application 18 / 930,681 entitled “POSITIONAL EMBEDDING AND TRAINING TECHNIQUES FOR A DIFFUSION MODEL”, filed on October 29, 2024, the entire contents of which are incorporated herein by reference.

[0131] like Figure 8 As further illustrated, the generative AI digital vision system 102 combines spatial embedding 812 and temporal embedding 810 to generate a spatial-temporal location code 818. In one or more embodiments, the spatial-temporal location code 818 refers to a data representation relating to both the spatial relationships and positions of visual elements within a frame, as well as motion information, action sequences, and transitions between frames (e.g., sequential dependencies between frames) within a frame sequence. Specifically, the spatial-temporal location code 818 includes a combined data representation capturing information from both visual and temporal dimensions. Accordingly, the generative AI digital vision system 102 utilizes the spatial-temporal location code 818 to remove noise from noise markers (e.g., the noise set of image markers 808) in a high-quality and accurate manner (e.g., by incorporating the context indicated by the data into the spatial-temporal location code 818).

[0132] Furthermore, such as Figure 8 As shown, the generative AI digital vision system 102 combines / adds image tags 814 with noise 816 to a spatial-temporal location code 818 (e.g., to generate a noise tag with a combination of spatial-temporal location codes). Therefore, the generative AI digital vision system 102 uses a diffusion-transformer model to process the image tags 814 with noise 816 and the spatial-temporal location code 818 to remove noise based on the spatial-temporal location code 818. Again, the generative AI digital vision system 102 utilizes the spatial-temporal location code 818 concatenated with the set of anchor tags (discussed above) as a guide for removing noise from the noise tags.

[0133] In one or more embodiments, the generative AI digital vision system 102 utilizes spatial-temporal location coding 818 (e.g., noise markers having a combination of spatial-temporal location coding) as the basis for anchoring image markers(s) or entire frames in the generated digital video. Specifically, spatial-temporal location coding 818 is a combined data representation that captures information from both the visual and temporal dimensions (e.g., the coding contains information such as where visual components should be located in space and time).

[0134] Thus, the generative AI digital vision system 102 uses spatial-temporal location coding 818 as an anchor to instruct the diffusion converter model to include image tags, multiple image tags, or entire image frames at specific instances in the generated digital video. For example, the generative AI digital vision system 102 utilizes spatial-temporal location coding 818 to instruct that image blocks (e.g., image blocks corresponding to image tags) should be included in every other frame of the digital video. In other words, the generative AI digital vision system 102 uses spatial-temporal location coding 818 to instruct the spatial and temporal inclusion of image tags in the generated digital video.

[0135] To further illustrate, the generative AI digital vision system 102 receives visual cues (e.g., digital images) and further receives instructions indicating that the generated digital video should include a specific object (e.g., a car) depicted in a digital image at the upper right corner of each frame of the generated digital video. In response, the generative AI digital vision system 102 generates a spatial-temporal location code 818 indicating spatial location (e.g., upper right corner) and temporal location (e.g., each frame of the digital video) by adjusting the associated time step, so that the spatial-temporal location code 818 is fully denoised and therefore not denoised during the denoising process.

[0136] In some instances, the generative AI digital vision system 102 receives visual cues and also receives instructions that the generated digital video should include the entire visual cue (e.g., a digital image) every 3 seconds within the generated digital video (e.g., at the 25-second mark of the digital video). In response, the generative AI digital vision system 102 generates a spatial-temporal location code 818 to conform to the received instructions.

[0137] although Figure 8 While the generation of spatial-temporal location codes 818 within the training context has been discussed, in one or more embodiments, the generative AI digital vision system 102 also uses spatial-temporal location codes 818 during inference. Specifically, in response to an image-to-video request from a client device, the generative AI digital vision system 102 generates spatial-temporal location codes for the digital image portion of a visual cue. Furthermore, the generative AI digital vision system 102 generates spatial-temporal location codes for any media attributes indicated by the client device.

[0138] For example, media attributes include the type of media (e.g., image or video), the format of the media, the subject matter of the media, the style of the media, the mood or thematic content, and any additional details (e.g., aspect ratio, frames per second, lens size, camera angle, type of motion such as zooming in or out, etc.).

[0139] Figure 9 The illustration shows a generative AI digital vision system 102 that generates a loss metric and modifies the parameters of a diffusion converter model according to one or more embodiments. (As described above...) Figure 7 As discussed herein, the generative AI digital vision system 102 generates image labels for training digital videos, adds noise to the image labels (e.g., a training label set, and the generative AI digital vision system 102 does not add noise to the training labels of anchor / adjustment frames), and removes noise from the noise labels (e.g., noisy training labels). Specifically, Figure 9 Anchor tag set 902 (e.g., for adjusting / anchor frames, also referred to as training anchor tag set), noise tag 904, and noise tag 906 are shown.

[0140] also, Figure 9 A generative AI digital vision system 102 is illustrated, which utilizes a diffusion converter model 908 (e.g., which includes converter block 910) to process anchor tag set 902, noise tag 904, and noise tag 906. Furthermore, Figure 9 A generative AI digital vision system 102 is shown that uses a diffusion converter model 908 to generate denoised markers 912.

[0141] In one or more embodiments, the generative AI digital vision system 102 uses a diffusion converter model 908 (e.g., a single-stream converter) to generate denoised labels 912 (e.g., denoised training labels) from noisy labels. Specifically, the denoised label 912 refers to a clean version of the data with noise removed from the labels. For example, over multiple denoising time steps (e.g., converter blocks), the generative AI digital vision system 102 utilizes the diffusion converter model 908 to remove noise from the noisy labels based on the anchor label set 902.

[0142] As further shown, the generative AI digital vision system 102 also uses a de-labeled model 914 (e.g., mentioned above). Figure 6 The de-labeling model 602 described herein is used to generate a denoised embedding 916 (e.g., a denoised training embedding). As shown, the generative AI digital vision system 102 compares the denoised embedding 916 with an embedding 918. Specifically, the embedding 918 originates from the generative AI digital vision system 102, which initially utilizes an encoder to generate embeddings from a sequence of frames in a training digital video (e.g., a de-labeling model 602 described herein). Figure 7 (As shown). In other words, the generative AI digital vision system 102 compares the pre-labeled form of the frame sequence of the training digital video with the denoised embedding 916 to determine the accuracy level of the diffusion converter model 908.

[0143] like Figure 9As shown, the generative AI digital vision system 102 generates a loss metric 920 by comparing a denoised embedding 916 and an embedding 918. In one or more embodiments, the generative AI digital vision system 102 determines the loss metric 920 by comparing the similarity between the predicted embedding (e.g., the denoised embedding) and the ground truth embedding. Specifically, the generative AI digital vision system 102 determines a mean squared error (MSE) loss to measure the mean squared error between corresponding elements of the predicted embedding and the ground truth embedding. For example, the goal of the MSE loss is to minimize the error between the prediction and the ground truth.

[0144] Go to Figure 10 Additional details regarding the various components and capabilities of the generative AI digital vision system 102 will now be provided. Specifically, Figure 10 The illustration shows an example schematic diagram of components 1000 to 1012 of a computing device 1000 (e.g., multiple servers 104 and / or client devices 110) implementing a generative AI digital vision system 102 according to one or more embodiments of the present disclosure. Figure 10 As illustrated, the generative AI digital vision system 102 includes a frame anchor system 105, a diffusion converter model manager 1002, an image-to-video request manager, an anchor tag manager 1006, a digital video manager 1010, and a storage manager 1012.

[0145] The diffusion converter model manager 1002 generates denoised tags. For example, the diffusion converter model manager 1002 utilizes a streamlined architecture without additional modulation or conditioning layers to process the input in a single-stream manner. Specifically, the diffusion converter model manager 1002 manages the training and optimization of the diffusion converter model. For example, the diffusion converter model manager 1002 receives training digital video, generates various embeddings / tags, and further generates loss metrics to modify the parameters of the diffusion converter model.

[0146] Image-to-video request manager 1004 receives one or more media requests from a client device. For example, image-to-video request manager 1004 provides a graphical user interface to the client device for inputting data for the image-to-video request. In one or more embodiments, image-to-video request manager 1004 provides the client device with options to upload one or more digital images, input text describing unconditional and / or conditional prompts, and further allows the client device to adjust preset parameters (e.g., digital video parameters such as camera angle, lighting, speed, frame rate, etc.). Additionally, in one or more embodiments, image-to-video request manager 1004 passes this data to additional components.

[0147] Anchor tag manager 1006 generates an anchor tag set. For example, anchor tag manager 1006 receives a digital image from image-to-video request manager 1004 and further determines that no noise should be added to the digital image. Further, anchor tag manager 1006 generates an embedding from the received digital image and further tokenizes the embedding (e.g., generating an image tag set). Additionally, anchor tag manager 1006 adds a time-step embedding to the image tag set to indicate that the image tag set is completely denoised. In doing so, anchor tag manager 1006 instructs diffusion converter model manager 1002 to use the anchor tag set as a guide for removing noise from the noise tags.

[0148] The denoising tag manager 1008 generates denoised tags. For example, the denoising tag manager 1008 uses a diffusion converter model to process the anchor tag set and the noise tags. In addition, the denoising tag manager 1008 utilizes multiple converter blocks of the diffusion converter model to remove noise from the noise tags to generate a denoised tag set.

[0149] The digital video manager 1010 generates digital video. For example, the digital video manager 1010 generates digital video from denoised tags. Specifically, the digital video manager 1010 de-tags the denoised tags and further utilizes a decoder to create digital video from the embeddings (e.g., denoised embeddings). Furthermore, the digital video manager 1010 enables the graphical user interface of a client device to display the generated digital video.

[0150] Storage manager 1012 stores various components generated by generative AI digital vision system 102. For example, storage manager 1012 stores model parameters (e.g., initial parameters and modified parameters) of diffusion converter model, cues (e.g., visual and textual), digital videos generated in response to cues, anchor tags, noise tags, training digital videos, embeddings, loss metrics, and additional training / startup data used to prepare diffusion converter model to generate digital videos from digital images.

[0151] Each of components 1002 to 1012 of the generative AI digital vision system 102 may include software, hardware, or both. For example, components 1002 to 1012 may include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (such as client devices or server devices). When executed by one or more processors, the computer-executable instructions of the generative AI digital vision system 102 may cause the computing devices(s) to perform the methods described herein. Alternatively, components 1002 to 1012 may include hardware, such as a dedicated processing device that performs a function or group of functions. Alternatively, components 1002 to 1012 of the generative AI digital vision system 102 may include a combination of computer-executable instructions and hardware.

[0152] Furthermore, components 1002 to 1012 of the generative AI digital vision system 102 can be implemented, for example, as one or more operating systems, one or more standalone applications, one or more modules of an application, one or more plugins, one or more library functions or functions that can be called by other applications, and / or cloud computing models. Therefore, components 1002 to 1012 of the generative AI digital vision system 102 can be implemented as standalone applications, such as desktop or mobile applications. Additionally, components 1002 to 1012 of the generative AI digital vision system 102 can be implemented as one or more web-based applications hosted on a remote server. Alternatively or additionally, components 1002 to 1012 of the generative AI digital vision system 102 can be implemented in a suite of mobile device applications or an "app". For example, in one or more embodiments, the generative AI digital vision system 102 may include digital software applications (such as...) FIREFLY AFTER EFFECTS CC PREMIERER RUSH, and / or It can be operated using PREMIERE PRO CC or in conjunction with digital software applications.

[0153] Figures 1 to 10 The corresponding text and examples provide various methods, systems, devices, and non-transitory computer-readable media from 1002 to 1012. In addition to the foregoing, one or more embodiments may also be described based on flowcharts including actions for achieving a particular result, such as... Figure 11 As shown. Figure 11It can be performed with more or fewer actions. Furthermore, these actions can be performed in different orders. Additionally, the actions described herein can be repeated or performed in parallel with each other or with different instances of the same or similar actions.

[0154] Figure 11 The illustration shows a flowchart of a series of actions 1100 for generating digital video based on denoising markers according to one or more embodiments. Figure 11 The illustration shows actions according to one embodiment; alternative embodiments may omit, add, reorder, and / or modify them. Figure 11 Any of the actions shown. In some implementations, Figure 11 The action is performed as part of the method. For example, in one or more embodiments, Figure 11 The actions are performed as part of a computer-implemented method. Alternatively, a non-transitory computer-readable medium may store instructions thereon that, when executed by at least one processor, cause the computing device to perform... Figure 11 The system performs the action. In one or more embodiments, the system executes... Figure 11 The system performs actions. For example, in one or more embodiments, the system includes at least one memory device. The system also includes at least one server device configured to cause the system to perform... Figure 11 The action.

[0155] Action series 1100 includes action 1102, which generates a set of image tags from a digital image. Further, action series 1100 includes action 1104, which generates a set of anchor tags from the set of image tags. Further, action series 1100 includes action 1105, which generates combined tags from the set of anchor tags and noise tags. Additionally, action series 1100 includes action 1106, which generates denoised tags from the noise tags and the set of anchor tags. Further, action series 1100 includes action 1108, which generates digital video based on the denoised tags.

[0156] Specifically, action 1102 includes generating a set of image tags from the digital image from an image-to-video request. Further, action 1104 includes generating a set of anchor tags from the set of image tags by adding a time-step embedding to the set of image tags, the time-step embedding indicating that the set of image tags is fully denoised. Further, action 1105 includes generating combined tags from the set of anchor tags and noise tags generated from noise. Furthermore, action 1106 includes processing the combined tags using a diffusion converter model to generate denoised tags. Furthermore, action 1108 includes generating a digital video comprising at least a portion of the digital image based on the denoised tags.

[0157] For example, in one or more embodiments, the action series 1100 includes receiving an image-to-video request from a client device during an inference period to generate a digital video comprising a digital image, and the image-to-video request instructing that the digital image be depicted in the digital video. Furthermore, in one or more embodiments, the action series 1100 includes wherein the image-to-video request instructs that the digital image be included as a first frame of a frame sequence, an intermediate frame of a frame sequence, or a final frame of a frame sequence.

[0158] Furthermore, in one or more embodiments, the series of actions 1100 includes generating an embedding of the digital image representing the image from an image-to-video request. Further, in one or more embodiments, the series of actions 1100 includes generating a set of image tags from the embedding by decomposing the digital image into multiple image chunks using a tokenization model.

[0159] Furthermore, in one or more embodiments, action series 1100 includes initializing noise tags by sampling random levels of noise from a noise distribution. Further, in one or more embodiments, action series 1100 includes removing noise from the noise tags based on a set of anchor tags using a diffusion converter model.

[0160] Furthermore, in one or more embodiments, motion series 1100 includes generating a frame sequence from denoised markers, wherein the digital video includes digital images as part of frames in the frame sequence, one or more keyframes in the frame sequence, or at least one of one or more motion frames in the frame sequence.

[0161] Additionally, in one or more embodiments, the action series 1100 includes generating a training embedding representing a frame from a training frame sequence. Furthermore, in one or more embodiments, the action series 1100 includes generating a set of training tags from the training embeddings using a tokenization model. Further, in one or more embodiments, the action series 1100 includes generating noisy training tags from a training frame sequence that does not include the frame using a tokenization model.

[0162] Furthermore, in one or more embodiments, action series 1100 includes generating training anchor tags by embedding time steps into a training tag set to instruct the training tag set to be fully denoised. Additionally, in one or more embodiments, action series 1100 includes utilizing a diffusion-transformer model to process the training anchor tags and noisy training tags to generate denoised training tags.

[0163] Furthermore, in one or more embodiments, action series 1100 includes generating denoised training embeddings from denoised training tags using a de-labeling model. Further, in one or more embodiments, action series 1100 includes comparing the denoised training embeddings with embeddings generated from a sequence of training frames prior to labeling.

[0164] Furthermore, in one or more embodiments, action series 1100 includes determining a loss metric by comparing the denoised training embeddings with embeddings generated from a sequence of training frames prior to tokenization to modify the parameters of the diffusion converter model. Furthermore, in one or more embodiments, action series 1100 includes generating a conditional labeled output from the denoised tags from a first process of combined tags passed through the diffusion converter model. Further, in one or more embodiments, action series 1100 includes generating an unconditional labeled output from the additional denoised tags from a second process of additional combined tags passed through the diffusion converter model. Furthermore, in one or more embodiments, action series 1100 includes generating a final labeled output by combining the conditional labeled output and the unconditional labeled output. Further, in one or more embodiments, action series 1100 includes generating a digital video comprising at least a portion of a digital image based on the final labeled output.

[0165] Further, in one or more embodiments, the action series 1100 includes generating an image tag set from a digital image as part of an image-to-video request. Additionally, in one or more embodiments, the action series 1100 includes generating a conditional tag output from a first process of combined tags passed through a trained diffusion converter model, wherein the combined tags include an anchor tag set from the image tag set and noise tags. Further, in one or more embodiments, the action series 1100 includes generating an unconditional tag output from a second process of additional combined tags passed through a trained diffusion converter model, wherein the additional combined tags include an anchor tag set and additional noise tags. Furthermore, in one or more embodiments, the action series 1100 includes generating a final tag output by combining the conditional tag output and the unconditional tag output. Furthermore, in one or more embodiments, the action series 1100 includes generating a digital video comprising at least a portion of the digital image based on the final tag output.

[0166] Furthermore, in one or more embodiments, the action series 1100 includes receiving an image to video request from a client device during an inference period, and a text prompt indicating that the digital image will be depicted in the digital video.

[0167] Furthermore, in one or more embodiments, the action series 1100 includes receiving an image-to-video request from a client device, the image-to-video request including a conditional cue for including a trained diffusion converter model in digital video and an unconditional cue for the digital video. Additionally, in one or more embodiments, the action series 1100 includes a conditional cue indicating that the digital image will be depicted as at least one of the initial frame, intermediate frame, or subset of frames in the digital video.

[0168] Furthermore, in one or more embodiments, the action series 1100 includes generating text tags based on conditional prompts for text prompts in an image-to-video request. Further, in one or more embodiments, the action series 1100 includes generating an image embedding from a digital image. Furthermore, in one or more embodiments, the action series 1100 includes initializing noise tags from a noise distribution. Further, in one or more embodiments, the action series 1100 includes combining text tags, image embeddings, and noise tags.

[0169] Furthermore, in one or more embodiments, the series of actions 1100 includes generating an image tag set from the image embedding by decomposing a digital image into multiple image blocks using a tagging model, thereby generating an anchor tag set from the image embedding. Further, in one or more embodiments, the series of actions 1100 includes adding a time-step embedding to the image tag set from the image embedding, wherein the time-step embedding instructs a trained diffusion converter model that the image tag set has been fully denoised.

[0170] Furthermore, in one or more embodiments, the action series 1100 includes generating text tags based on unconditional prompts for text cues in an image-to-video request. Further, in one or more embodiments, the action series 1100 includes generating an image embedding from a digital image. Furthermore, in one or more embodiments, the action series 1100 includes initializing noise tags from a noise distribution. Further, in one or more embodiments, the action series 1100 includes combining text tags, image embeddings, and noise tags.

[0171] Furthermore, in one or more embodiments, the action series 1100 includes interpolating between conditionally labeled outputs and unconditionally labeled outputs using a classifier-free guided model, wherein the interpolation includes guided scaling of the digital video generated by a trained diffusion converter model based on the conditionally labeled outputs. Further, in one or more embodiments, the action series 1100 includes generating a final labeled output based on the interpolation from the classifier-free guided model.

[0172] Furthermore, in one or more embodiments, the action series 1100 includes receiving an image-to-video request from a client device during an inference period to generate a digital video including the digital image. Further, in one or more embodiments, the action series 1100 includes a scenario where the image-to-video request indicates that the digital image will be included as part of a frame in a frame sequence, one or more keyframes in the frame sequence, or at least one of one or more motion frames in the frame sequence.

[0173] Furthermore, in one or more embodiments, the action series 1100 includes training a diffusion converter model by adding noise to a subset of frames in a frame sequence of a training video, wherein the subset of frames does not include one or more frames that serve as anchor frames in the training video.

[0174] Further, in one or more embodiments, action series 1100 includes generating a conditional tag output from denoised tags from a first process of combined tags via a diffusion converter model, wherein the combined tags include a set of anchor tags and denoised tags. Additionally, in one or more embodiments, action series 1100 includes generating an unconditional tag output from additional denoised tags from additional combined tags via a diffusion converter model, wherein the additional combined tags include a set of anchor tags and additional denoised tags. Further, in one or more embodiments, action series 1100 includes generating a final tag output by combining the conditional tag output and the unconditional tag output. In one or more embodiments, action series 1100 includes generating a digital video comprising at least a portion of a digital image based on the final tag output.

[0175] Figure 12 Examples of diffusion models 1200 according to various aspects of this disclosure are shown. In some examples, diffusion model 1200 describes a reference... Figure 14 The operation and architecture of the described diffusion converter model (e.g., the single-stream diffusion converter model described above). Figure 12 The diffusion model depicted is an example of or includes aspects of the generative AI digital vision system 102 as described herein. Specifically, the diffusion-transformer model combines the principles of the diffusion model with those of the transformer model. Accordingly, Figure 12 A generative AI digital vision system 102 is illustrated, which initializes a trained diffusion converter model by utilizing a forward diffusion process to corrupt data and then using a denoising process to create media from the corrupted data. In other words, the generative AI digital vision system 102 teaches a diffusion converter model to create generative content from noise using a forward diffusion process and a denoising process.

[0176] As an example, the diffusion model is a generative model that operates by progressively corrupting / noising the input signal and learning from the corrupted data to generate new samples. Specifically, the diffusion model uses a forward diffusion process to add noise at a series of time steps and a backward diffusion process to remove noise at multiple time steps corresponding to the number of forward steps.

[0177] As a further example, converter models are designed for sequence-to-sequence modeling tasks (e.g., tasks that generate data based on sequence order, such as generative tasks utilizing labels). As mentioned above, converter models typically include self-attention mechanisms and multilayer perceptrons (e.g., feedforward networks that move data through a converter architecture). Furthermore, converter architectures often consider location information, such as the spatial-temporal location encoding discussed above.

[0178] Within the context of the diffusion and converter models discussed above, this paper provides additional details on how the diffusion-converter model combines the principles of diffusion and converter models. A diffusion-converter model is a class of generative neural networks that can be trained to generate new data with features similar to those found in the training data. Specifically, diffusion-converter models can be used to generate novel media items such as images, audio files, videos, 3D models, or other digital media items. Diffusion-converter models can be used for a variety of media processing tasks, including image super-resolution, generating media items with perceptual metrics, image inpainting, and media manipulation. Specifically, the diffusion-converter model differs from existing diffusion model architectures in that it combines a converter architecture with diffusion principles for removing noise from noise tags. Specifically, the architecture of the diffusion-converter model in this disclosure includes self-attention layers and multilayer perceptrons. In one or more embodiments, instead of conditioning the input, the diffusion-converter model includes positional encodings and other (clean) tags (such as anchor tags) along with the noise tags as guidance on how the converter block should remove noise from the noise tags.

[0179] In one or more embodiments, the diffusion-transformer model leverages the architecture of a transformer model to capture long-range dependencies and complex structures in high-dimensional data. Specifically, the diffusion-transformer model operates by processing labeled data of images and text to fully account for long-range dependencies. Furthermore, the diffusion-transformer model uses a transformer architecture to predict denoised data at each time step (e.g., transformer block) and, as described above, uses a self-attention mechanism on noisy data to understand how noise should be removed across various noisy input labels.

[0180] As discussed in detail above, the generative AI digital vision system 102 utilizes a diffusion transformer model (instead of the UNet diffusion architecture), where the generative AI digital vision system 102 uses an encoder (e.g., a VAE encoder) to abstract pixel details into latent representations (e.g., embeddings). For example, the generative AI digital vision system 102 uses the encoder to abstract pixel data into semantic information that can be adapted for use in the transformer architecture (e.g., the transformer architecture captures global context from the latent representation through attention).

[0181] Furthermore, in one or more embodiments, the generative AI digital vision system 102 does not inject diffusion information via adaLN modulation, but instead designs a diffusion converter model in a single-stream manner. In other words, the generative AI digital vision system 102 utilizes a diffusion converter model with input inflow and input outflow in a single stream. Therefore, in one or more embodiments, the generative AI digital vision system 102 does not utilize adaLN modulation for input conditioning and feeds position encoding, anchor tags, and / or other encoded information (e.g., tag-level diffusion time step embedding) directly into the self-attention layer along with noise tags.

[0182] In one or more embodiments, methods for operating diffusion models include Denoising Diffusion Probabilistic Model (DDPM) and Denoising Diffusion Implicit Model (DDIM). In DDPM, the generative process includes an inverted random Markov diffusion process. On the other hand, DDIM uses a deterministic process such that the same input produces the same output. In some cases, DDIM can reduce the number of time steps during media generation. The diffusion model can also be characterized by whether noise is added to the media item itself or to the media features generated by the encoder (i.e., latent diffusion). In the pixel diffusion model, noise is added and removed in the pixel space. In the latent diffusion model, noise is added (and removed) in the latent space of the media features rather than in the pixel space. Thus, the latent diffusion model uses back-diffusion to generate media features, and these media features can be decoded to obtain synthetic media items.

[0183] In one or more embodiments, the generative AI digital vision system 102 utilizes a diffusion process to add noise to the media data 1205, which has already been transformed from pixel space 1210 to latent space. For example, the generative AI digital vision system 102 transforms data 1205 to latent space, adds noise to data 1205, and then uses a converter block to denoise the noisy data 1220 (e.g., remove noise from noise markers to obtain a synthesized media item). Specifically, Figure 12 Data 1220 (e.g., visual cues) processed by encoder 1212 (e.g., to generate embeddings and then further generate tags) is shown.

[0184] also, Figure 12 A generative AI digital vision system 102 is shown that adds noise to data 1205 using a forward diffusion process 1215. Furthermore, Figure 12 A generative AI digital vision system 102 is shown that removes noise from noisy data 1220 using a denoising process 1225. For example, Figure 12A generative AI digital vision system 102 is illustrated, which generates media 1230 using decoder 1229. Further, in one or more embodiments, the generative AI digital vision system 102 adds noise to the data in a gradual manner (e.g., over multiple time steps corresponding to multiple converter blocks). In doing so, the generative AI digital vision system 102 trains a diffusion converter model to create generative content from the corrupted data (e.g., noisy data 1220).

[0185] As just mentioned, Figure 12 Data 1220 processed by encoder 1212 is shown. As discussed in some details above, the diffusion converter model operates by processing labeled data. To this end, the generative AI digital vision system 102 uses encoder 1212 to generate embeddings(s) of data 1220 and further converts data 1220 into labels. For example, the transformation process is referred to as labeling or chunking. Specifically, labeling / chunking starts with an image having dimensions of HxWxC, where H is the height, W is the width, and C is the number of color channels (3 for RGB images).

[0186] In one or more embodiments, the generative AI digital vision system 102 uses chunking to divide an image into chunks (e.g., partitioning the image into non-overlapping chunks of size PxP), where each chunk contains PxPxC pixel values. Further, for an image of size HxW, the total number of chunks would be (H / P)x(W / P). Additionally, the generative AI digital vision system 102 flattens the image chunks into one-dimensional vectors to create a chunk embedding sequence similar to text tags (e.g., in the context of natural language processing).

[0187] Furthermore, the generative AI digital vision system 102 generates image labels by performing linear projection to convert image patches into fixed-dimensional embeddings (d-dimensional vectors). Specifically, the generated image labels serve as input labels for a diffusion converter model.

[0188] Figure 13 Examples of a method 1300 for media generation according to various aspects of this disclosure are shown. In some examples, method 1300 describes the operation of a diffusion converter model, such as referencing... Figure 12 The application of the described diffusion model 1200. In some examples, these operations are performed by a system that includes a processor that executes a set of code to control functional elements such as the generative AI digital vision system 102 described above.

[0189] Alternatively or additionally, the steps of method 1300 may be performed using dedicated hardware. Typically, these operations are performed according to the methods and processes described in various aspects of this disclosure. In some cases, the operations described herein are performed as a series of sub-steps or in combination with other operations.

[0190] At operation 1305, the user provides textual and / or visual cues describing the content to be included in the generated media item. For example, the user can provide the cue "a person playing with a cat." In some examples, guidance can be provided in forms other than text, such as via images (e.g., visual cues), sketches, audio input, or layout.

[0191] At operation 1310, the system converts the text prompt (or other prompt guidance) into a tokenized or other multidimensional representation compatible with a single stream diffusion converter model. For example, a converter model or a multimodal encoder can be used to transform the text into a vector or a series of vectors. In some cases, the encoder for generating the tokens is trained independently of the diffusion model (e.g., via a trained dual-VAE model, which is described in DUAL-VAE FOR MORE EFFICIENT AND EFFECTIVE DIFFUSIONMODEL TRAINING and incorporated above by reference).

[0192] At operation 1315, a noise map including random noise is initialized. The noise map can be in pixel space or latent space. By utilizing random noise to initialize media items, different variations of media items including content described by the cues can be generated. At operation 1320, the system generates media items based on the noise map, tags from the cues (e.g., text cues and / or visual cues), and additional spatial-temporal location encoding.

[0193] Figure 14 The diffusion process 1400 according to various aspects of this disclosure is illustrated. Specifically, Figure 14 Additional details on the operational principles of the diffusion model are provided. Accordingly, Figure 14 Context and details are provided regarding the principles of borrowing a self-diffusion model to aid in operating the diffusion converter model. In some examples, diffusion process 1400 describes the operation of the diffusion converter model, such as referencing... Figure 12 The denoising process of the diffusion model 1200 is described.

[0194] As referenced above Figure 12The described use of a diffusion converter model can involve a process for initializing noise (e.g., generating noise markers in the latent space) and a denoising process 1410 for denoising the noise markers to obtain denoised markers. The denoising process 1410 can be represented as p(x) t-1 |x t In some cases, the neural network is trained to perform the denoising process 1410 (i.e., to remove noise continuously).

[0195] In the example forward pass of the latent diffusion model, the model uses a Markov chain to map the observed variable x0 (the embedding in the latent space) to intermediate variables x1,…,x T When latent variables are passed through a neural network such as a diffusion-transformer model, the Markov chain progressively adds Gaussian noise to the data (e.g., embeddings such as visual signals) to obtain an approximate posterior q(x). 1:T |x0), where x1,…,x T It has the same dimensions as x0.

[0196] A neural network can be trained to perform denoising. During the denoising process 1410, the model uses noisy data x T (Such as noise labeling) begins and the data is denoised to obtain p(x) t-1 |x t In each step t-1, the denoising process 1410 employs x. t Examples of such elements include first intermediate denoising markers, spatial-temporal location encoding, and markers (e.g., indicating cues). Here, t represents a converter block in a sequence of converter blocks associated with different noise levels. The denoising process 1410 iteratively outputs x. t-1 Such as the second intermediate denoising marker, up to x T Return to the fully denoised marker x0. The denoising process can be represented as:

[0197] p θ (x t-1 |x t :=N(x t-1 μ θ (x t ,t),Σ θ (x t ,t)). (1)

[0198] Furthermore, the process of adding noise to data to generate noise labels is expressed as the joint probability of the sample sequences in a Markov chain, which can be written as the product of conditional probability and marginal probability:

[0199]

[0200] Where p(x)T )=N(x T The distribution (0, I) is pure noise because the reverse process uses the result of the forward process and pure noise samples as input, and This represents the Gaussian transition sequence corresponding to the sequence to which Gaussian noise is added to the sample.

[0201] During inference, the observed data x0 in the pixel space can be mapped to the latent space as input, and the generated data can be... The output is mapped back from the latent space to the pixel space (e.g., using a decoder in a trained dual VAE model). In some examples, x0 represents the original clean label, and the latent variables x1,…,x T Indicates noise markers, and This indicates the generated high-quality item.

[0202] Figure 15 This is a flowchart of a step-by-step process 1500 in an example implementation of the algorithm, depicting it as an operation for training a machine learning model. In one or more embodiments, process 1500 describes the operations described for configuring the training components of the diffusion converter model. Process 1500 provides one or more examples of generating training data, using the training data to train a machine learning model, and using the trained machine learning model to perform a task.

[0203] At the beginning of this example, the machine learning system collects training data (box 1502), which will be used as the basis for training the machine learning model; that is, the training data defines what is being modeled. Training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, service provider system platforms that reveal application programming interfaces (e.g., social media platforms), user data collection systems (e.g., digital surveys and online crowdsourcing systems), and so on. Training data collection may also include data augmentation and synthetic data generation techniques to expand and diversify the available training data, balancing techniques to balance multiple positive and negative samples, and so on.

[0204] Machine learning systems can also be configured to identify features relevant to the type of task the machine learning model is being trained on (box 1504). Examples of tasks include classification, natural language processing, generative artificial intelligence, recommendation engines, reinforcement learning, clustering, etc. To this end, the machine learning system collects training data based on the identified features and / or filters the training data based on the identified features after collection. The training data is then used to train the machine learning model.

[0205] To train the machine learning model in the illustrated example, the machine learning model is first initialized (box 1506). Initializing the machine learning model includes selecting the model architecture to be trained (box 1508). Examples of model architectures include neural networks, diffusion-transformer models, transformer models, diffusion models, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning neural networks, etc.

[0206] A loss function is also selected (box 1510). The loss function is used to measure the difference between the output (i.e., the prediction) of the machine learning model and the target value used to train the machine learning model (e.g., as expressed by the training data). Additionally, an optimization algorithm 1512 is selected, which is used in conjunction with the loss function to optimize the parameters of the machine learning model during training; examples include gradient descent, stochastic gradient descent (SGD), etc.

[0207] Initializing a machine learning model also includes setting initial values ​​for the model (box 1514) (box 1516). Examples of machine learning models include initializing node weights and biases as part of training to improve the efficiency of training and computational resource consumption. Hyperparameters are also set to control the training of the machine learning model; examples include regularization parameters, model parameters (e.g., the number of layers in a neural network), learning rate, batch size of the training data, etc. Various techniques are used to set hyperparameters, including using randomization techniques, using heuristics learned from other training scenarios, etc.

[0208] The machine learning system then uses the training data to train a machine learning model (Box 1518). A machine learning model refers to a computer representation that can be adjusted (e.g., trained and retrained) based on inputs of training data to approximate an unknown function. Specifically, the term machine learning model can include models that learn from known data and make predictions about known data by using algorithms (e.g., using the model architecture described above) to learn and relearn by analyzing training data to generate outputs that reflect the patterns and properties expressed by the training data.

[0209] Examples of training types include supervised learning using labeled data, unsupervised learning involving finding underlying structures or patterns within the training data, reinforcement learning based on optimization functions (e.g., rewards and / or penalties), and using nodes as part of “deep learning.” For example, a machine learning model can be configured to include multiple nodes that collectively form multiple layers. These layers can be configured to include an input layer, an output layer, and one or more hidden layers. Computation is performed by the nodes within a layer through the hidden states, via a system of weighted connections “learned” during training, for example, by optimizing the performance of the machine learning model to perform the associated task using a chosen loss function and backpropagation.

[0210] As part of training the machine learning model, it is determined whether a stopping criterion (decision box 1520) is met, i.e., to validate the machine learning model. This stopping criterion can be used to reduce overfitting of the machine learning model, reduce computational resource consumption, and improve the machine learning model's ability to handle previously unseen data (i.e., data not specifically included as examples in the training data). Examples of stopping criteria include, but are not limited to, a predetermined number of periods, validating loss stability, achieving a performance improvement threshold, whether a threshold level of accuracy is met, or performance metrics such as accuracy and recall. If the stopping criterion has not yet been met ("No" from decision box 1520), then in this example, process 1500 continues training the machine learning model using the training data (box 1518).

[0211] If the stopping criteria ("Yes" from decision box 1520) are met, the trained machine learning model is used to generate an output based on subsequent data (box 1522). For example, the trained machine learning model is trained to perform the task as described above, and therefore, once trained, is configured to perform the task based on subsequent data received as input and processed by the machine learning model.

[0212] Embodiments of this disclosure may include or utilize a dedicated or general-purpose computer including computer hardware, such as one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more processes described herein may be implemented at least in part as instructions executed by one or more computing devices (e.g., any of the media content access devices described herein). Typically, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory) and executes those instructions to perform one or more processes, including one or more processes described herein.

[0213] Computer-readable media can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of this disclosure may include at least two distinctly different types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0214] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drive (“SSD”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.

[0215] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately regards that connection as a transport medium. A transport medium may include networks and / or data links, which may be used to carry desired program code in the form of computer-executable instructions or data structures, and which may be accessible by general-purpose or special-purpose computers. Combinations of the above should also be included within the scope of computer-readable media.

[0216] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be cached in RAM within a network interface module (e.g., a "NIC") and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Therefore, it should be understood that a non-transitory computer-readable storage medium (device) can be included in computer system components that also (or even primarily) utilize the transmission medium.

[0217] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or group of functions. In one or more embodiments, the computer-executable instructions are executed on a general-purpose computer to convert the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. The computer-executable instructions may be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter as defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the described features and actions are disclosed as exemplary forms of implementing the claims.

[0218] Those skilled in the art will appreciate that this disclosure can be practiced in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframes, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. This disclosure can also be practiced in distributed system environments, where local and remote computer systems linked via a network (via hardwired data links, wireless data links, or a combination of hardwired and wireless data links) perform tasks. In a distributed system environment, program modules can reside in local and remote memory storage devices.

[0219] The embodiments of this disclosure can also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly configured via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0220] Cloud computing models can be composed of various characteristics, such as on-demand self-service, broadband network access, resource pooling, rapid elasticity, and metric services. Cloud computing models can also reveal various service models, such as Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“IaaS”). Different deployment models (such as private cloud, community cloud, public cloud, hybrid cloud, etc.) can also be used to deploy cloud computing models. In this specification and claims, a “cloud computing environment” means an environment employing cloud computing.

[0221] Figure 16Examples of a method 1600 for training a diffusion model according to various aspects of this disclosure are shown. In some embodiments, method 1600 is described as referenced... Figure 18 The described operations are for configuring the training components of the diffusion converter model. Method 1600 represents the operations used for training as referenced above. Figure 14 Examples of the described reverse diffusion process. In some examples, these operations are performed by a system including a processor that executes a set of code to control the functional elements of the device, such as... Figure 12 The guided diffusion model described in [the text].

[0222] Additionally or alternatively, certain procedures of method 1600 may be performed using dedicated hardware. Typically, these operations are performed according to the methods and procedures described in various aspects of this disclosure. In some cases, the operations described herein are performed as a series of sub-steps or in combination with other operations.

[0223] At operation 1605, the user initializes the untrained model. Initialization may include defining the model's architecture and establishing initial values ​​for the model parameters. In some cases, initialization may include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, and the location of hop connections.

[0224] At operation 1610, the system uses a forward diffusion process to add noise to the medium term in N stages. In some cases, the forward diffusion process is a fixed process, in which Gaussian noise is continuously added to the medium term. In latent diffusion models (e.g., labeled spaces), Gaussian noise can be continuously added to features in the latent space.

[0225] At operation 1615, starting from stage N, the system uses a backdiffusion process at each stage n to predict the output or features of stage n-1. For example, the backdiffusion process can predict noise added by the forward diffusion process and remove the predicted noise from the noise input to obtain the predicted output. In some cases, the original media terms are predicted at each stage of the training process.

[0226] At operation 1620, the system compares the predicted output (or feature) at stage n-1 with the actual media item (or feature) (such as the output at stage n-1 or the original input). For example, given observed data x, a diffusion model can be trained to minimize the negative log-likelihood of the training data -logp. θ The variational upper limit of (x).

[0227] At operation 1625, the system updates the model's parameters based on this comparison. For example, gradient descent can be used to update the U-Net's parameters. Time-dependent parameters of the Gaussian transition can also be learned. However, in some embodiments, for the diffusion converter model, the system uses mean squared error denoising loss to update the parameters of each converter block.

[0228] Figure 17 Examples of computing devices 1700 according to various aspects of this disclosure are shown. The computing device 1700 may be an example of a generative AI digital media system device (e.g., a device for interacting with a generative AI digital vision system 102, as described above). In one aspect, the computing device 1700 includes processor(s) 1705, a memory subsystem 1710, a communication interface 1715, an I / O interface 1720, user interface(s) 1725, and a channel 1730.

[0229] In one or more embodiments, computing device 1700 is an example of or includes aspects of the generative AI digital vision system 102 described above. In one or more embodiments, computing device 1700 includes one or more processors 1705 capable of executing instructions stored in memory subsystem 1710 to perform media generation.

[0230] According to some aspects, computing device 1700 includes one or more processors 1705. In some cases, the processor is an intelligent hardware device (e.g., a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof). In some cases, the processor is configured to use a memory controller to operate a memory array. In other cases, the memory controller is integrated into the processor. In some cases, the processor is configured to execute computer-readable instructions stored in memory to perform various functions. In one or more embodiments, the processor includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0231] According to some aspects, the memory subsystem 1710 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software containing instructions that, when executed, cause the processor to perform the various functions described herein. In some cases, among others, the memory contains a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller operates the memory cells. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within the memory store information in the form of logical states.

[0232] According to some aspects, communication interface 1715 operates at the boundary between communication entities (such as computing device 1700, one or more user devices, a cloud, and one or more databases) and channel 1730, and can record and process communications. In some cases, communication interface 1715 is provided to implement a processing system coupled to a transceiver (e.g., a transmitter and / or receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for communication devices via an antenna.

[0233] Depending on the context, I / O interface 1720 is controlled by an I / O controller to manage the input and output signals of computing device 1700. In some cases, I / O interface 1720 manages peripheral devices not integrated into computing device 1700. In some cases, I / O interface 1720 represents a physical connection or port to an external peripheral device. In some cases, the I / O controller uses, for example... The operating system or other known operating systems. In some cases, the I / O controller represents a modem, keyboard, mouse, touchscreen, or similar device, or the I / O controller interacts with them. In some cases, the I / O controller is implemented as a component of the processor. In some cases, the user interacts with the device via the I / O interface 1720 or via hardware components controlled by the I / O controller.

[0234] According to some aspects, the user interface component(s) 1725 enables a user to interact with the computing device 1700. In some cases, the user interface component(s) 1725 includes audio devices such as an external speaker system, external display devices such as a display screen, input devices (e.g., remote control devices that interface directly with the user interface or via an I / O controller), or combinations thereof. In some cases, the user interface component(s) 1725 includes a GUI.

[0235] Figure 18 An example of a generative AI digital media system apparatus 1800 according to various aspects of this disclosure is shown. The generative AI digital media system apparatus 1800 may include reference... Figure 12 Examples of the described diffusion model or aspects thereof. In one or more embodiments, the generative AI digital media system device 1800 includes a processor unit 1805, a memory unit 1810, a diffusion converter model 1815, an I / O module 1820, and a training component 1825. The training component 1825 updates the parameters of the diffusion converter model 1815 stored in the memory unit 1810. In some examples, the training component 1825 is located outside the generative AI digital media system device 1800.

[0236] Processor unit 1805 includes one or more processors. A processor is an intelligent hardware device, such as a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or any combination thereof.

[0237] In some cases, processor unit 1805 is configured to use a memory controller to operate a memory array. In other cases, the memory controller is integrated into processor unit 1805. In some cases, processor unit 1805 is configured to execute computer-readable instructions stored in memory unit 1810 to perform various functions. In some aspects, processor unit 1805 includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing. According to some aspects, processor unit 1805 includes reference... Figure 17 One or more processors as described.

[0238] Memory unit 1810 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause at least one processor of processor unit 1805 to perform the various functions described herein.

[0239] In some cases, memory cell 1810 includes a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, memory cell 1810 includes a memory controller that operates the memory cells within memory cell 1810. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within memory cell 1810 store information in the form of logical states. According to some aspects, memory cell 1810 is a reference... Figure 17 An example of the described memory subsystem 1710.

[0240] According to some aspects, the generative AI digital media system apparatus 1800 uses one or more processors of the processor unit 1805 to execute instructions stored in the memory unit 1810 to perform the functions described herein. For example, the generative AI digital media system apparatus performs the operations described in the following aspects.

[0241] Memory unit 1810 may include a diffusion converter model 1815, which is trained to remove noise from noise tags based on spatial-temporal location encoding. For example, after training, the diffusion converter model 1815 may perform a reference... Figures 12 to 13 The described inference operation is used to remove noise from noise markers and generate media such as videos and / or images.

[0242] In one or more embodiments, the diffusion converter model 1815 is an artificial neural network (ANN). An ANN can be a hardware or software component comprising connection nodes (i.e., artificial neurons) loosely corresponding to neurons in the human brain. Each connection or edge transmits a signal from one node to another (like a physical synapse in the brain). When a node receives a signal, it processes the signal and then transmits the processed signal to other connection nodes. Specifically, each converter block of the diffusion converter model 1815 can represent a connection node.

[0243] ANNs have multiple parameters, including weights and biases associated with each neuron in the network. These parameters control the degree of connection between neurons and affect the neural network's ability to capture complex patterns in data. These parameters, also known as model parameters or model weights, are variables that determine the behavior and characteristics of a machine learning model. Accordingly, a multilayer perceptron within each converter block of a diffusion-transformer model represents aspects of an ANN.

[0244] In some cases, the signals between nodes consist of real numbers, and the output of each node is computed as a function of its inputs. For example, nodes can use other mathematical algorithms to determine their outputs, such as selecting the maximum value from the input as the output, or any other suitable algorithm used to activate the node. Each node and edge is associated with one or more node weights that determine how signals are processed and transmitted. In some cases, nodes have a threshold below which no signal is transmitted at all. In some examples, nodes are aggregated into layers.

[0245] The parameters of the diffusion-transformer model 1815 can be organized into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer and the last layer is called the output layer. In some cases, the signal passes through certain layers multiple times. Hidden (or intermediate) layers contain hidden nodes and are located between the input and output layers. Hidden layers perform nonlinear transformations on the inputs fed into the network. Each hidden layer is trained to produce a defined output that contributes to the joint output of the ANN's output layers. The hidden representation is a machine-readable data representation of the inputs learned from the ANN's hidden layers and produced by the output layers. As the understanding of the ANN's inputs improves as the ANN is trained, the hidden representations gradually differentiate themselves from earlier iterations.

[0246] Training component 1825 can train diffusion converter model 1815. For example, the parameters of diffusion converter model 1815 can be learned or estimated from training data and then used to make predictions or perform tasks based on learned patterns and relationships in the data. In some examples, parameters are tuned during the training process to minimize a loss function or maximize a performance metric. The goal of the training process may be to find optimal values ​​of the parameters that allow the machine learning model to make accurate predictions or perform well on a given task.

[0247] Correspondingly, node weights can be adjusted to improve the accuracy of the output (i.e., by minimizing the loss corresponding in some way to the difference between the current result and the target result). Edge weights increase or decrease the strength of the signal transmitted between nodes. For example, during training, the algorithm adjusts machine learning parameters according to optimization techniques such as gradient descent, stochastic gradient descent, or other optimization algorithms to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the diffusion converter model 1815 can be used to make predictions on new, unseen data (i.e., during inference).

[0248] I / O module 1820 receives input from generative AI digital media system device 1800 and transmits its output to other devices or users. For example, I / O module 1820 receives input to diffusion converter model 1815 and transmits the output of diffusion converter model 1815. According to some aspects, I / O module 1820 is a reference... Figure 17 An example of the described I / O interface 1720.

Claims

1. A computer-implemented method, comprising: From image-to-video requests that include digital images, generate a set of image tags from the digital images; An anchor tag set is generated from the image tag set by adding a time step embedding, wherein the time step embedding indicates that the anchor tag set is completely denoised; Generate a combined tag from the set of anchor tags and the noise tags generated from the noise; The combined tags are processed using a diffusion converter model to generate denoised tags; as well as A digital video comprising at least a portion of the digital image is generated based on the denoising markers.

2. The computer-implemented method according to claim 1 further includes: During the inference period, an image-to-video request is received from the client device to generate the digital video including the digital image, and the image-to-video request indicates that the digital image should be depicted in the digital video. The image-to-video request indicates that the digital image will be included as the first frame of the frame sequence, the middle frame of the frame sequence, or the last frame of the frame sequence.

3. The computer-implemented method according to claim 1, wherein generating the image tag set comprises: From the image to the video request, the digital image generates an embedding representing the digital image; as well as The digital image is decomposed into multiple image blocks using a tokenization model to generate the image tag set from the embedding.

4. The computer-implemented method according to claim 1, wherein generating the denoised markers comprises: The noise tag is initialized by sampling noise at random levels from the noise distribution; as well as Using the diffusion converter model, noise is removed from the noise markers based on the set of anchor markers.

5. The computer-implemented method of claim 1, wherein generating the digital video comprises generating a frame sequence from the denoised markers, wherein the digital video comprises the digital image as at least one of the following: a portion of frames in the frame sequence, one or more keyframes in the frame sequence, or one or more motion frames in the frame sequence.

6. The computer-implemented method of claim 1, further comprising training the diffusion converter model by: Generate training embeddings representing the frames from the training frame sequence; A training label set is generated from the training embeddings using a labeling model; and The tokenization model is used to generate noisy training tags from the training frame sequence that does not include the frame.

7. The computer-implemented method according to claim 6, further comprising: Training anchor tags are generated by embedding time steps into the training tag set to indicate that the training tag set is completely denoised. as well as The training anchor tags and the noise training tags are processed using the diffusion converter model to generate denoised training tags.

8. The computer-implemented method according to claim 7, further comprising: A de-labeling model is used to generate denoised training embeddings from the denoised training labels; as well as Compare the denoised training embedding with the embedding generated from the training frame sequence before tokenization; as well as The loss metric is determined by comparing the denoised training embedding with the embedding generated from the training frame sequence before tokenization to modify the parameters of the diffusion converter model.

9. The computer-implemented method according to claim 1, wherein generating the digital video comprises: A conditional tag output from the denoised tag is generated from the first flow of the combined tag through the diffusion converter model; An unconditional tag output from the additional denoised tags is generated from a second process of additional combined tags through the diffusion converter model; The final tag output is generated by combining the conditional tag output and the unconditional tag output; as well as The digital video, comprising at least a portion of the digital image, is generated based on the final labeled output.

10. A system comprising: Memory components; as well as One or more processing devices coupled to the memory component, the one or more processing devices being configured to perform operations including: Generate a set of image tags from digital images that are part of an image-to-video request; Conditional label output is generated from a first process of combined labels obtained through a trained diffusion converter model, wherein the combined labels include a set of anchor labels and noise labels from the image label set; An unconditional label output is generated from a second process of additional combined labels through the trained diffusion converter model, wherein the additional combined labels include the anchor label set and additional noise labels; The final tag output is generated by combining the conditional tag output and the unconditional tag output; and A digital video comprising at least a portion of the digital image is generated based on the final labeled output.

11. The system of claim 10, wherein the operation includes receiving, during an inference period, an image-to-video request from a client device that generates the digital video and a text prompt indicating that the digital image will be depicted in the digital video.

12. The system of claim 10, wherein the operation includes: The image-to-video request is received from the client device, the image-to-video request including conditional cues for the trained diffusion converter model to be included in the digital video, and unconditional cues for the digital video. The conditional prompt indicates that the digital image will be depicted as at least one of the following: the initial frame, intermediate frame, or subset of frames in the digital video.

13. The system of claim 10, wherein generating the combined tag comprises: Generate text tags from the conditional prompts in the text prompts for the image to video request; Generate an image embedding from the digital image; Initialize the noise tag from the noise distribution; as well as Combine the text markers, the image embeddings, and the noise markers.

14. The system of claim 13, further comprising generating the anchor tag set from the image embedding by: The digital image is decomposed into multiple image blocks using a tokenization model to generate the image tag set from the image embedding; and A time-step embedding is added to the set of image tags from the image embedding, wherein the time-step embedding instructs the trained diffusion converter model that the set of image tags is completely denoised.

15. The system of claim 10, wherein generating the additional combined tag comprises: Generate text tags from the unconditional prompts in the text prompts for the image to video request; Generate an image embedding from the digital image; Initialize the noise tag from the noise distribution; as well as Combine the text markers, the image embeddings, and the noise markers.

16. The system of claim 10, wherein combining the conditional flag output and the unconditional flag output comprises: The classifier-free guided model interpolates between the conditionally labeled output and the unconditionally labeled output, wherein the interpolation includes guided scaling, which facilitates the trained diffusion converter model to generate the digital video based on the conditionally labeled output; as well as The final labeled output is generated based on the interpolation of the classifier-free guided model.

17. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform an operation, the operation comprising: From image-to-video requests that include digital images, generate a set of image tags from the digital images; An anchor tag set is generated from the image tag set by adding a time step embedding, wherein the time step embedding indicates that the image tag set is completely denoised; Generate a combined tag from the set of anchor tags and the noise tags generated from the noise; The combined tags are processed using a diffusion converter model to generate denoised tags; as well as A digital video comprising at least a portion of the digital image is generated based on the denoising markers.

18. The non-transitory computer-readable medium of claim 17, wherein the operation further comprises: During the inference period, the image-to-video request that generates the digital video including the digital image is received from the client device. The image-to-video request indicates that the digital image will be included as at least one of the following: a portion of a frame in a frame sequence, one or more keyframes in the frame sequence, or one or more motion frames in the frame sequence.

19. The non-transitory computer-readable medium of claim 17, wherein the operation further comprises training the diffusion converter model by adding noise to a subset of frames of a frame sequence of a training video, wherein the subset of frames does not include one or more frames that serve as anchor frames in the training video.

20. The non-transitory computer-readable medium of claim 17, wherein generating the digital video comprises: A conditional tag output from the denoised tags is generated from the first process of the combined tags through the diffusion converter model, wherein the combined tags include the anchor tag set and the noise tags; An unconditional tag output from additional denoising tags is generated from a second process of additional combined tags through the diffusion converter model, wherein the additional combined tags include the anchor tag set and additional noise tags; The final tag output is generated by combining the conditional tag output and the unconditional tag output; as well as The digital video, comprising at least a portion of the digital image, is generated based on the final labeled output.

Citation Information

Patent Citations

  • Dual-VAE for more efficient and effective diffusion model training

    US20260073579A1

  • Positional embedding and training techniques for a diffusion model

    US20260073692A1