Grid diffusion model device and method for text-to-video generation

The grid diffusion model addresses the challenges of high computational costs and large data sets in text-to-video generation by reducing video dimension to image dimension, enabling efficient and high-quality video production with fixed GPU memory and small data sets.

WO2025206474A1PCT designated stage Publication Date: 2025-10-02UNIST (ULSAN NAT INST OF SCI & TECH)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/009985
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2024-07-12
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Generating high-quality videos from text is challenging due to the larger data sets and higher computational costs compared to image generation, and existing methods require complex architectures and large-scale text-to-video pair datasets.

Method used

A grid diffusion model that extracts a fixed number of grid images identified by text, reducing the video dimension to the image dimension, using autoregressive grid image interpolation to generate high-quality videos with a fixed amount of GPU memory, regardless of the number of frames.

Benefits of technology

Enables efficient high-quality video generation with reduced GPU memory costs and small training data sets, allowing for various image-based methods to be applied to videos, maintaining temporal consistency and generating more frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024009985_02102025_PF_FP_ABST
    Figure KR2024009985_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A grid diffusion model method and device for text-to-video generation are disclosed. The grid diffusion model method for text-to-video generation, according to one embodiment of the present invention, may comprise the steps of: generating a key grid image corresponding to text; and using an interpolation model that performs autoregressive grid image interpolation, thereby creating, from the key grid image, an interpolation grid image constituting video.
Need to check novelty before this filing date? Find Prior Art

Description

Grid diffusion model device and method for generating video from text

[0001] The present invention relates to a grid diffusion model device and method for generating video from text, which generates a plurality of grid images having a time flow according to a prompt input by a user, and creates a generative video through a unique interpolation process using each of the grid images.

[0002] Registration No. 10-2649818 (March 18, 2024) "3D lip-sync video generation device and method"

[0003] Registration No. 10-2649818 discloses a configuration for generating a 3D lip-sync video of a 3D model of a person in a 2D video speaking based on speech audio unrelated to the 2D video.

[0004] Registration No. 10-2613143 (December 8, 2023) "Text-to-video generation method, device, equipment, and medium"

[0005] Registration number 10-2613143 discloses a configuration that receives text input from a user in response to a user touch detected in an area where the initial screen is located, and generates a video based on the text input so that it can be posted to an information sharing app.

[0006] Publication No. 10-2023-0095432 (June 29, 2023) "Text-based Character Animation Synthesis System"

[0007] Publication number 10-2023-0095432 discloses a configuration that learns a character animation generation model from a large amount of video and text and motion information extracted therefrom, and synthesizes character animation corresponding to new text information corresponding to fingerprints or dialogue based on the model.

[0008] Recently, text-to-image generation technology for generating images from text has been greatly improved due to the development of diffusion models.

[0009] Furthermore, text-to-image generation technology is evolving into text-to-video generation technology.

[0010] However, generating videos from text is a more difficult task than generating images from text due to much larger data sets and higher computational costs.

[0011] Most existing video generation methods use 3D U-Net architectures or autoregressive generation methods that consider the temporal dimension.

[0012] These methods require large data sets and may be computationally limited compared to text-to-image generation.

[0013] Therefore, there is an urgent need for a novel, simple and effective grid diffusion technique for text-to-video generation without temporal dimension from architectures and large-scale text-to-video pair datasets.

[0014] An embodiment of the present invention aims to provide a grid diffusion model device and method for video generation from text, which extracts a fixed number of grid images identified by text and uses these grid images to generate high-quality video using a fixed amount of GPU memory regardless of the number of frames.

[0015] Furthermore, embodiments of the present invention aim to apply various image-based methods to videos, such as image manipulation and text-guided video manipulation, by reducing the dimension of a video to the dimension of an image.

[0016] Additionally, embodiments of the present invention aim to demonstrate the suitability of the model for actual video generation.

[0017] According to one embodiment of the present invention, a grid diffusion model method for generating video from text may include the steps of: generating a key grid image corresponding to the text; and creating an interpolated grid image constituting a video from the key grid image through an interpolation model that performs autoregressive grid image interpolation.

[0018] In addition, a grid diffusion model device for generating video from text according to an embodiment of the present invention may be configured to include a generation unit that generates a key grid image corresponding to the text; and a processing unit that creates an interpolation grid image constituting a video from the key grid image through an interpolation model that performs autoregressive grid image interpolation.

[0019] According to one embodiment of the present invention, a grid diffusion model device and method for video generation from text can be provided, which extracts a fixed number of grid images identified by text, and generates a high-quality video using these grid images using a fixed amount of GPU memory regardless of the number of frames.

[0020] Additionally, according to one embodiment of the present invention, since the dimension of a video is reduced to the dimension of an image, various image-based methods can be applied to videos, such as image manipulation and text-guided video manipulation.

[0021] Additionally, according to one embodiment of the present invention, the suitability of the model for actual video generation can be proven.

[0022] FIG. 1 is a block diagram illustrating the configuration of a grid diffusion model device for generating video from text according to one embodiment of the present invention.

[0023] Figure 2 is a diagram to visually demonstrate the training of the key grid image generation model.

[0024] Figure 3 is a diagram for explaining the operation of the grid diffusion model device of the present invention.

[0025] Figure 4 is a diagram illustrating an example of training an interpolation model.

[0026] Figure 5 is a diagram illustrating the process of generating a generative video.

[0027] FIG. 6 is a diagram illustrating an example of video manipulation using an image manipulation technique according to the present invention.

[0028] FIG. 7 is a flowchart illustrating a grid diffusion model method for generating video from text according to one embodiment of the present invention.

[0029] Hereinafter, embodiments are described in detail with reference to the attached drawings. However, the embodiments may be modified in various ways, and the scope of the patent application is not limited or restricted by these embodiments. It should be understood that all modifications, equivalents, or alternatives to the embodiments are included within the scope of the patent application.

[0030] The terms used in the examples are for illustrative purposes only and should not be construed as limiting. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "comprise" or "have" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood to not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the embodiments pertain. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0032] In addition, when describing with reference to the attached drawings, identical components will be assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted. When describing embodiments, if a detailed description of a related known technology is judged to unnecessarily obscure the gist of the embodiment, the detailed description will be omitted.

[0033] FIG. 1 is a block diagram illustrating the configuration of a grid diffusion model device for generating video from text according to one embodiment of the present invention.

[0034] Referring to FIG. 1, a grid diffusion model device (hereinafter, abbreviated as 'grid diffusion model device') (100) for generating video from text according to an embodiment of the present invention may be configured to include a generation unit (110), a processing unit (120), and a model training unit (130). In addition, the grid diffusion model device (100) may be configured to include a key grid image generation model (115) and an interpolation model (131) composed of 1-Step and 2-Step interpolation models (132, 133).

[0035] First, the generation unit (110) generates a key grid image corresponding to the text. That is, the generation unit (110) can extract grid images whose contents match or whose conditions are satisfied by the input text from the original video, and collect these grid images to generate a key grid image.

[0036] In generating a key grid image, the generation unit (110) can receive a prompt as the text.

[0037] Here, a prompt refers to a place where input is received from a user or where a program instructs a user what action to perform, and in the present invention, it may refer to a command tool that receives text input from a user for interaction with the user.

[0038] The generation unit (110) can output m internal frames (where m is a natural number greater than or equal to 4) specified by the prompt from the key grid image generation model (115).

[0039] For example, when the prompt 'Teddy bear dancing disco in the starry night' is input, the generation unit (110) can learn the prompt from the key grid image generation model (115) and output four grid images f1, f10, f19, and f28, each having the appearance of a teddy bear dancing disco in the starry night, as internal frames from the key grid image generation model (115).

[0040] Additionally, the generation unit (110) can generate the key grid image by arranging the m internal frames in chronological order.

[0041] In the example described above, the generation unit (110) can generate a key grid image by arranging the four output grid images f1, f10, f19, and f28 in quadrants in the order of the time at which each grid image was captured. The generation unit (110) can generate a key grid image by arranging the grid image f1, which is the earliest in time among the four grid images, to the first quadrant, and the grid image f28, which is the latest in time, to the fourth quadrant.

[0042] The above key grid image generation model (115) can perform dimension reduction by selecting an internal frame of an image dimension from a video frame of a video dimension. That is, the key grid image generation model (115) has a function of dimensionally reducing a processing target from a video to an image by selecting and outputting a specific internal frame from an original video frame according to a prompt.

[0043] The processing unit (120) creates an interpolated grid image that constitutes a video from the key grid image through an interpolation model (131) that performs autoregressive grid image interpolation. That is, the processing unit (120) can play a role in interpolating the internal frames within the key grid image into an interpolated grid image that fills in the gaps between the internal frames by having the interpolation model (131) learn the internal frames within the key grid image.

[0044] Here, the interpolation model (131) can be configured to include a 1-Step interpolation model (132) and a 2-Step interpolation model (133).

[0045] The 1-Step interpolation model (132) may be a model that outputs a 1-Step interpolation grid image by an interpolation operation using an internal frame within a key grid image.

[0046] The 2-Step interpolation model (133) may be a model that outputs a 2-Step interpolation grid image by an interpolation operation using the first image frame within a 1-Step interpolation grid image.

[0047] The model training unit (130) can play a role in training the interpolation model (131) in advance.

[0048] In training the interpolation model (131), the model training unit (130) can first select four grid images (f1 to f4) prior to the reference time point (t5) from a plurality of grid images (f1, ... ft-1, ft) as previous grid images. That is, the model training unit (130) can select four grid images f1 to f4 captured at times t1 to t4 prior to the reference time point (t5) as previous grid images. At this time, the previous grid images can be arranged from the first quadrant to the fourth quadrant in the order of capture time of each grid image f1 to f4.

[0049] In addition, the model training unit (130) can create a masked grid image by blanking any two grid images among the four grid images (f5 to f8) after the reference time point (t5) and including the remaining grid images. That is, the model training unit (130) can create a masked grid image by blanking the grid images f6 and f7, which were captured at intermediate times among the four grid images f5 to f8 captured at times t5 to t8 after the reference time point (t5). Accordingly, the masked grid image can have grid images f5 and f8 arranged in the first and fourth quadrants, respectively, and blanks arranged in the second and third quadrants.

[0050] Thereafter, the model training unit (130) can train the interpolation model (131) to fill in the blanks in the masked grid image based on the previous grid image. That is, the model training unit (130) can train the interpolation model (131) to generate grid images to be inserted into the blanks by autoregressive grid image interpolation utilizing the previous grid image.

[0051] The processing unit (120) can create an interpolation grid image from a key grid image by utilizing a pre-trained interpolation model (131).

[0052] Specifically, the processing unit (120) can create a first mask grid image using the internal frame included in the key grid image, and then create a 1-Step interpolation grid image by learning from a 1-Step interpolation model (132) among the interpolation models.

[0053] The processing unit (120) can create the first mask grid image by placing the internal frames included in the key grid image in the first and fourth quadrants and placing the blanks in the second and third quadrants.

[0054] In addition, the processing unit (120) can create the 1-Step interpolation grid image by interpolating so that the blanks placed in the 2nd and 3rd quadrants of the first mask grid image are filled by the 1-Step interpolation model (132).

[0055] For the above-described key grid image example having four grid images f1, f10, f19, and f28 as internal frames, regarding a teddy bear, the processing unit (120) can select grid images f1 and f10 and create a first mask grid image (M(1)) that places f1 in the first quadrant, f10 in the fourth quadrant, and blanks in the second and third quadrants.

[0056] Thereafter, the processing unit (120) can create a 1-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the first mask grid image (M(1)) with grid images f4 and f7 by learning the created first mask grid image (M(1)) in the 1-Step interpolation model (132).

[0057] In addition, the processing unit (120) can create a 2-Step interpolation grid image by using the first image frame included in the 1-Step interpolation grid image to create n second mask grid images (where n is a natural number greater than or equal to 3), and then learning from the 2-Step interpolation model (133) among the interpolation models (131).

[0058] The processing unit (120) can create the n second mask grid images by considering a combination of arranging each of the first image frames in the 1st and 4th quadrants and arranging the blanks in the 2nd and 3rd quadrants.

[0059] In addition, the processing unit (120) can create the 2-Step interpolation grid image by interpolating so that the blanks arranged in the 2nd and 3rd quadrants of the n second mask grid images are filled by the 2-Step interpolation model (133).

[0060] For the above-described 1-Step interpolation grid image having the four grid images f1, f4, f7, and f10 as the first image frame, the processing unit (120) can create a second mask grid image (M(2-1)) that places f1 in the first quadrant, f4 in the fourth quadrant, and blanks in the second and third quadrants, a second mask grid image (M(2-2)) that places f4 in the first quadrant, f7 in the fourth quadrant, and blanks in the second and third quadrants, and a second mask grid image (M(2-3)) that places f7 in the first quadrant, f10 in the fourth quadrant, and blanks in the second and third quadrants.

[0061] Thereafter, the processing unit (120) can create a 2-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the 2nd mask grid image (M(2-1)) with grid images f2 and f3 by learning the 2-Step interpolation model (133) of the 2nd mask grid image (M(2-1)).

[0062] In addition, the processing unit (120) can create a 2-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the 2nd mask grid image (M(2-2)) with grid images f5 and f6 by learning the 2-Step interpolation model (133) of the 2nd mask grid image (M(2-2)).

[0063] Finally, the processing unit (120) can create a 2-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the 2nd mask grid image (M(2-3)) with grid images f8 and f9 by learning the 2-Step interpolation model (133) of the 2nd mask grid image (M(2-3)).

[0064] In general, the processing unit (120) can create a 2-Step interpolation grid image by interpolating between f1 and f10 with grid images f2 to f9 through learning in the 1-Step interpolation model (132) and the 2-Step interpolation model (133).

[0065] Accordingly, the processing unit (120) can output a generated video by sequentially connecting the second image frames included in the 2-Step interpolation grid image.

[0066] That is, the processing unit (120) can output a video of natural movement by connecting grid images f1 to f10, which are second image frames included in a 2-Step interpolation grid image.

[0067] In addition, the processing unit (120) may continuously output the generative video after the grid image f10 by repeating the above-described process (creating the first mask grid image (M(1)), learning in the 1-Step interpolation model (132) and the 2-Step interpolation model (133), etc.) by sequentially selecting other internal frames of the key grid image regarding the teddy bear (e.g., selecting f10 and f19 and selecting f19 and f28).

[0068] At this time, the time interval between the first image frames included in the 1-Step interpolation grid image may be smaller than the time interval between the internal frames included in the key grid image, and may also be larger than the time interval between the second image frames included in the 2-Step interpolation grid image. That is, in the grid diffusion model device (100), the time interval between the second image frames included in the 2-Step interpolation grid image may be made as small as possible so that the grid images are continuously connected when outputting the generated video, thereby enabling natural video playback.

[0069] According to one embodiment of the present invention, a grid diffusion model device and method for video generation from text can be provided, which extracts a fixed number of grid images identified by text, and generates a high-quality video using these grid images using a fixed amount of GPU memory regardless of the number of frames.

[0070] Additionally, according to one embodiment of the present invention, since the dimension of a video is reduced to the dimension of an image, various image-based methods can be applied to videos, such as image manipulation and text-guided video manipulation.

[0071] Additionally, according to one embodiment of the present invention, the suitability of the model for actual video generation can be proven.

[0072] The development of diffusion models has greatly improved the performance of text-to-image models.

[0073] Unlike GAN-based models, diffusion models offer desirable properties such as a wide distribution, fixed training objectives, and easy scalability, making training easier.

[0074] In relation to the diffusion model, various studies are being conducted on manipulating or generating images from text, and research on generating videos from text is also being actively pursued.

[0075] However, video generation is more difficult than generating images from text because of the higher dimensionality of videos and the scarcity of text-to-video datasets, which incurs higher costs.

[0076] In previous studies, videos were generated by using additional temporal dimensions and super-resolution models to maintain temporal consistency and resolution of the videos.

[0077] These characteristics of video make efficiency an important issue in video generation, which is why many video generation studies focus on efficiency.

[0078] The grid diffusion model device (100) of the present invention, unlike existing video generation paradigms, reduces the high dimension of a video to the dimension of an image, thereby enabling high-quality video generation without significant GPU memory costs and large-scale pairing data sets.

[0079] The grid diffusion model device (100) can actively utilize the strength of diffusion to generate video from text.

[0080] The grid diffusion model device (100) can perform two steps: key grid image generation and autoregressive grid image interpolation.

[0081] Figure 2 is a diagram to visually demonstrate the training of the key grid image generation model.

[0082] As shown in FIG. 2, the grid diffusion model device (100) can express video frames of the video dimension, which are listed in chronological order, as grid images of the image dimension through dimension reduction.

[0083] In addition, the grid diffusion model device (100) can generate a key grid image as an output from the 2D U-Net by training the expressed grid image in the 2D U-Net, which is a key grid image generation model, according to a prompt.

[0084] In FIG. 2, the grid diffusion model device (100) can extract four grid images having the appearance of a child eating watermelon corresponding to the prompt 'Toddler eating ripe watermelon' through a 2D U-Net, and generate a Key Grid image including the four extracted grid images as an internal frame.

[0085] To reduce the video dimension to an image dimension, the grid diffusion model device (100) can sequentially select and arrange grid images of a specific time in the video to generate a key grid image.

[0086] A key grid image may consist of, for example, four inner frames, selected by text.

[0087] The grid diffusion model device (100) can fine-tune a pre-trained key grid image generation model using prompts as a condition for generating a key grid image.

[0088] The grid diffusion model device (100) can be generated as a key grid image including four internal frames having temporal consistency according to Stable Diffusion.

[0089] The grid diffusion model device (100) can interpolate internal frames of a key grid image while maintaining temporal consistency and order.

[0090] The grid diffusion model device (100) can be used by applying an image manipulation method because it reduces the video dimension to an image dimension.

[0091] The grid diffusion model device (100) can perform autoregressive grid image interpolation.

[0092] The grid diffusion model device (100) can configure an interpolation model that takes a masked grid image as input and uses a previously generated key grid image as a condition.

[0093] The grid diffusion model device (100) can connect the embedding spaces of two images in the latent dimension.

[0094] Through this, the grid diffusion model device (100) can generate a consistent video frame in which the current grid image and the previous grid image match.

[0095] Additionally, the grid diffusion model device (100) can generate the next key grid image by autoregressively using the previous key grid image to generate more frames.

[0096] Using this approach, the grid diffusion model device (100) can maintain temporal consistency and generate videos with more than 28 frames.

[0097] In addition, the grid diffusion model device (100) can be applied to various application fields such as image-based models for video manipulation by using image manipulation because it expresses video as a grid image.

[0098] The grid diffusion model device (100) achieves better performance than existing text-video models without a large paired training data set and can generate more frames at a fixed amount of GPU memory cost.

[0099] The grid diffusion model device (100) can apply grid diffusion for text-video generation to the real world.

[0100] The grid diffusion model device (100) can provide a simple yet effective new grid diffusion for efficient text-to-video generation by reducing the temporal dimension of the video.

[0101] The grid diffusion model device (100) can generate high-quality video using a fixed amount of GPU memory even with a small number of frames and a small amount of training data.

[0102] The grid diffusion model device (100) can easily apply an image-based model to corresponding video tasks such as video manipulation and video style editing because it expresses a video as a grid image.

[0103] The grid diffusion model device (100) can produce high-quality video that is faithful to text and exceeds standards.

[0104] Research on generating high-quality images from text (Text-to-Image Generation) has been ongoing for a long time, and recent advances in diffusion models have made it possible to generate high-quality images from plain text, which has had a significant social impact.

[0105] Recent studies have leveraged architectures such as Transformers, Variational Autoencoders (VAEs), and diffusion models to generate higher resolution and more general images from text descriptions.

[0106] For example, DALLE and Parti are training a transformer model on a large dataset of text-image pairs to be able to generate images from plain text inputs.

[0107] Meanwhile, models such as GLIDE, DALL2, and Stable Diffusion use diffusion models to generate images.

[0108] These diffusion-based models have shown promising results in image generation tasks.

[0109] The grid diffusion model device (100) provides an approach to generate high-quality video from text without temporal dimension by utilizing stable diffusion pre-trained on a large-scale text-image pair data set, leveraging the strengths of the diffusion model.

[0110] Text-to-video generation faces two major challenges: the lack of large-scale, high-quality text-to-video datasets and the complexity of modeling the temporal dimension.

[0111] NUWA introduces a unified representation space that enables effective multi-task learning for both text-to-image and text-to-video generation tasks.

[0112] CogVideo integrates an additional temporal attention module through multi-frame rate hierarchical training that ensures alignment between text and video, leveraging CogView-2, a pre-trained text-to-image generation model.

[0113] Make-A-Video is expanding into text-to-video by introducing a super-resolution strategy for high-quality, high-frame-rate video generation, leveraging DALL2, a diffusion-based text-to-image model.

[0114] The video diffusion model trains images and videos jointly with the 3D U-Net diffusion model architecture.

[0115] LatentShift generates videos by shifting spatial U-Net feature maps forward and backward in the temporal dimension, which ensures temporal consistency and efficiency in the video.

[0116] PYoCo extends the text-to-image diffusion model to 3D and fine-tunes a pre-trained diffusion model. PYoCo also utilizes a noise dictionary and a pre-trained eDiff-I model for video generation.

[0117] Despite active research, the field of text-to-video generation remains challenging due to complex model structures and the large amount of training data required.

[0118] The grid diffusion model device (100) can present a new paradigm for generating text-video without a large training data set of text-video pairs by solving these problems through a simple architecture with an effective approach.

[0119] The grid diffusion model device (100) proposes a simple yet effective new approach for text-to-video generation using a grid diffusion model.

[0120] Figure 3 is a diagram for explaining the operation of the grid diffusion model device of the present invention.

[0121] As illustrated in FIG. 3, the grid diffusion model device (100) can perform (a) key grid image generation and (b) autoregressive grid image interpolation.

[0122] (a) In the key grid image generation step, the grid diffusion model device (100) generates a key grid image representing a video from a given text.

[0123] In FIG. 3, for the prompt 'A hot air balloon is floating in the blue sky, and the camera is zoomed out at a high resolution of 4K.', the grid diffusion model device (100) can output four internal frames corresponding to a hot air balloon floating in the blue sky from the key grid image generation model and align them to generate a key grid image.

[0124] (b) In the autoregressive grid image interpolation step, the grid diffusion model device (100) can generate a generative video by interpolating the generated key grid image.

[0125] The grid diffusion model device (100) can create a mask grid image using an internal frame within a key grid image and then learn from a 1-Step interpolation model to create a 1-Step interpolation grid image.

[0126] In addition, the grid diffusion model device (100) can create a plurality of mask grid images using image frames within a 1-Step interpolation grid image and then learn from a 2-Step interpolation model to create a 2-Step interpolation grid image.

[0127] The grid diffusion model device (100) can generate a generative video by a combination of connecting image frames within a generated 2-Step interpolation grid image.

[0128] The grid diffusion model device (100) can generate high-quality video with a fixed amount of GPU memory cost and small training data, and can also manipulate the video at the image level.

[0129] The grid diffusion model device (100) can generate a key grid image for generating a video by reducing the temporal dimension.

[0130] A key grid image consists of four internal frames that represent key actions or events in the video.

[0131] The grid diffusion model device (100) can output four internal frames in time order through a key grid image generation model and generate a key grid image by arranging the four internal frames in the output order.

[0132] The grid diffusion model device (100) generates a key grid image having a resolution of 512×512, and each internal frame has a resolution of 254×254.

[0133] The grid diffusion model device (100) can train and fine-tune a grid image generation model.

[0134] The grid diffusion model device (100) can generate a key grid image that effectively expresses scene changes and dynamic movements from a key grid image generation model by appropriately reflecting the movement of the image data set.

[0135] The grid diffusion model device (100) can generate more image frames for video generation by using four internal frames within a key grid image.

[0136] Internal frames are interconnected, and temporal consistency between frames can be maintained.

[0137] The grid diffusion model device (100) can be configured to include an interpolation model (1-Step interpolation model, 2-Step interpolation model) that outputs a grid image by interpolating a parasitic grid image and a masked grid image in a self-regressive manner.

[0138] The grid diffusion model device (100) can train an interpolation model using grid images with different intervals for each interpolation model.

[0139] Figure 4 is a diagram illustrating an example of training an interpolation model.

[0140] As shown in FIG. 4, the grid diffusion model device (100) can train an interpolation model using the previous grid image and the masked grid image.

[0141] The grid diffusion model device (100) can generate a previous grid image of the image dimension by selecting grid images f1 to f4 captured at t1 to t4 before the reference time point t5 from among grid images (f1, ... Ft-1, ft) of the video dimension.

[0142] In addition, the grid diffusion model device (100) can generate a masked grid image of an image dimension in which the grid images f6 and f7 among the grid images f5 to f8 captured at t5 to t8 after the reference time point t5 are blanked and f5 and f8 are placed in the first and fourth quadrants.

[0143] The grid diffusion model device (100) can train an interpolation model so that blank grid images f6 and f7 within the masked grid image are filled in based on the previous grid image.

[0144] The grid diffusion model device (100) can train an interpolation model that fills a blank grid image, for example, through a command called “Fill in the blanks.”

[0145] Figure 5 is a diagram illustrating the process of generating a generative video.

[0146] As shown in FIG. 5, the grid diffusion model device (100) can generate a key grid image having four internal frames f1, f10, f19, and f28 spaced at nine frame intervals by utilizing a key grid image generation model.

[0147] The grid diffusion model device (100) can obtain a 1-Step interpolation grid image that interpolates f4 and f7 between f1 and f10 by creating a first mask grid image M(1) having f1, f10 and two blanks and then training the 1-Step interpolation model.

[0148] In addition, the grid diffusion model device (100) can obtain a 2-Step interpolation grid image including a plurality of temporally continuous image frames f2 to f9 between f1 and f10 by creating a second mask grid image M(2-1) having f1, f4 and two blanks, a second mask grid image M(2-2) having f4, f7 and two blanks, and a second mask grid image M(2-3) having f7, f10 and two blanks, and then training each of these in a 2-Step interpolation model.

[0149] The grid diffusion model device (100) can repeat the process of using a previously generated grid image as a conditioning image when interpolating between the first image frame and the fourth image frame in a grid image in an autoregressive manner to enhance temporal consistency.

[0150] The grid diffusion model device (100) can successfully generate a generative video that connects a total of 28 image frames (f1 to f28) by repeating the above process, filling in between f1, f28, and f1 and f28.

[0151] The grid diffusion model device (100) can train a key grid generation model to expand more image frames at a prompt.

[0152] The grid diffusion model device (100) can automatically regressively generate the next key grid image by utilizing the previous key grid image.

[0153] The grid diffusion model device (100) can interpolate internal frames within a newly generated key grid image.

[0154] The grid diffusion model device (100) can generate more image frames with both context and temporal consistency while complying with fixed GPU memory constraints.

[0155] The grid diffusion model device (100) is capable of video manipulation using text.

[0156] The grid diffusion model device (100) can apply various image-based methods to the video dimension by reducing the dimension of video generation to image generation.

[0157] The grid diffusion model device (100) can perform video manipulation according to text guides.

[0158] The grid diffusion model device (100) selects four internal frames from the original video according to an input prompt to generate a key grid image.

[0159] The grid diffusion model device (100) generates a generative video by interpolating a key grid image using an interpolation model.

[0160] The grid diffusion model device (100) can generate a video with temporal consistency by automatically regressing previously generated frames.

[0161] The grid diffusion model device (100) generates a higher resolution video.

[0162] Since the grid diffusion model device (100) generates video using a text-image model, a high-resolution text-image model can be applied to the grid diffusion method.

[0163] The grid diffusion model device (100) can generate an image with a resolution of 1024x1024.

[0164] The grid diffusion model device (100) can generate a video with a resolution of 510×510 by applying a 2×2 grid.

[0165] The grid diffusion model device (100) can flexibly expand high-resolution video generation using a text-image model.

[0166] The grid diffusion model device (100) proposes a new grid diffusion model for text-video generation to address the problems of large-scale text-video pairs, lack of data sets, and high GPU memory cost for video generation.

[0167] The grid diffusion model device (100) can generate high-quality video using a fixed amount of GPU memory regardless of the number of frames by expressing the video as a grid image.

[0168] The grid diffusion model device (100) can easily apply various image dimension methods to video manipulation.

[0169] FIG. 6 is a diagram illustrating an example of video manipulation using an image manipulation technique according to the present invention.

[0170] The grid diffusion model device (100) can easily manipulate video using text.

[0171] For image manipulation, the grid diffusion model device (100) can use a grid image created by selecting four frames from a Webvid-10M video as an input image.

[0172] A prompt is a set of conditions for manipulating a grid image.

[0173] Inter prompts are prompt conditions for interpolation models.

[0174] The grid diffusion model device (100) can edit content by inserting glasses or changing the shape of a hat, and can also change the style of the video.

[0175] As shown in Fig. 6, the grid diffusion model device (100) can select four internal frames in which a woman appears.

[0176] The grid diffusion model device (100) can output (generate) a key grid image from a key grid image generation model by manipulating images for four selected internal frames corresponding to an input prompt.

[0177] In Figure 6, a key grid image is illustrated that is generated by changing a girl to a man wearing a hat, corresponding to the prompt "Replace a girl to A man wearing a hat man".

[0178] Afterwards, the grid diffusion model device (100) can generate a final modified 28-frame video by learning the key grid image changed to a man wearing a hat from the 1-Step interpolation model and the 2-Step interpolation model described above.

[0179] Hereinafter, FIG. 7 describes in detail the work flow of the grid diffusion model device (100) according to embodiments of the present invention.

[0180] FIG. 7 is a flowchart illustrating a grid diffusion model method for generating video from text according to one embodiment of the present invention.

[0181] The grid diffusion model method for generating video from text according to the present embodiment can be performed by a grid diffusion model device (100).

[0182] First, the grid diffusion model device (100) generates a key grid image corresponding to the text (710). Step (710) may be a process of extracting grid images whose content matches or conditions are satisfied by the input text from the original video, and collecting these grid images to generate a key grid image.

[0183] In generating a key grid image, the grid diffusion model device (100) can receive a prompt as the above text.

[0184] Here, a prompt refers to a place where input is received from a user or where a program instructs a user what action to perform, and in the present invention, it may refer to a command tool that receives text input from a user for interaction with the user.

[0185] The grid diffusion model device (100) can output m internal frames (wherein m is a natural number greater than or equal to 4) specified by the prompt from the key grid image generation model.

[0186] For example, when the prompt 'Teddy bear dancing disco in the starry night' is input, the grid diffusion model device (100) can learn the prompt from the key grid image generation model and output four grid images f1, f10, f19, and f28, each having an image of a teddy bear dancing disco in the starry night, as internal frames from the key grid image generation model.

[0187] Additionally, the grid diffusion model device (100) can generate the key grid image by arranging the m internal frames in chronological order.

[0188] In the above-described example, the grid diffusion model device (100) can generate a key grid image by arranging the four output grid images f1, f10, f19, and f28 in quadrants in the order of the time at which each grid image was captured. The grid diffusion model device (100) can generate a key grid image by arranging the grid image f1, which is the earliest in time among the four grid images, to be placed in the first quadrant, and the grid image f28, which is the latest in time, to be placed in the fourth quadrant.

[0189] The above key grid image generation model can perform dimension reduction by selecting an internal frame of the image dimension from a video frame of the video dimension. In other words, the key grid image generation model has the function of dimensionally reducing the processing target from a video to an image by selecting and outputting a specific internal frame from the original video frame according to a prompt.

[0190] In addition, the grid diffusion model device (100) creates an interpolated grid image that constitutes a video from the key grid image through an interpolation model that performs autoregressive grid image interpolation (720). Step (720) may be a process of learning internal frames within the key grid image from the interpolation model and interpolating the internal frames into an interpolated grid image that fills in the gaps between the internal frames.

[0191] Here, the interpolation model can be configured to include a 1-Step interpolation model and a 2-Step interpolation model.

[0192] A 1-Step interpolation model may be a model that outputs a 1-Step interpolation grid image by interpolation operation using internal frames within a key grid image.

[0193] A 2-Step interpolation model may be a model that outputs a 2-Step interpolation grid image by an interpolation operation using the first image frame within a 1-Step interpolation grid image.

[0194] The grid diffusion model device (100) can train an interpolation model in advance.

[0195] In training the interpolation model (131), the grid diffusion model device (100) can first select four grid images (f1 to f4) prior to the reference time point (t5) from a plurality of grid images (f1, ... ft-1, ft) as previous grid images. That is, the grid diffusion model device (100) can select four grid images f1 to f4 captured at times t1 to t4 prior to the reference time point (t5) as previous grid images. At this time, the previous grid images can be arranged from the first quadrant to the fourth quadrant in the order of capture time of each grid image f1 to f4.

[0196] In addition, the grid diffusion model device (100) can create a masked grid image by blanking any two grid images among the four grid images (f5 to f8) after the reference time point (t5) and including the remaining grid images. That is, the grid diffusion model device (100) can create a masked grid image by blanking the grid images f6 and f7, which were captured at intermediate times, among the four grid images f5 to f8 captured at times t5 to t8 after the reference time point (t5). Accordingly, in the masked grid image, the grid images f5 and f8 can be arranged in the first and fourth quadrants, respectively, and blanks can be arranged in the second and third quadrants.

[0197] Thereafter, the grid diffusion model device (100) can train the interpolation model (131) to fill in the blanks in the masked grid image based on the previous grid image. That is, the grid diffusion model device (100) can train the interpolation model (131) to generate a grid image to be inserted into the blanks by autoregressive grid image interpolation utilizing the previous grid image.

[0198] The grid diffusion model device (100) can create an interpolated grid image from a key grid image by utilizing a pre-trained interpolation model (131).

[0199] Specifically, the grid diffusion model device (100) can create a 1-Step interpolation grid image by creating a first mask grid image using an internal frame included in the key grid image, and then learning from a 1-Step interpolation model (132) among the interpolation models.

[0200] The grid diffusion model device (100) can create the first mask grid image by placing internal frames included in the key grid image in quadrants 1 and 4 and placing blanks in quadrants 2 and 3.

[0201] In addition, the grid diffusion model device (100) can create the 1-Step interpolation grid image by interpolating so that the blanks arranged in the 2nd and 3rd quadrants of the first mask grid image are filled by the 1-Step interpolation model (132).

[0202] For the above-described key grid image example having four grid images f1, f10, f19, f28 as internal frames, regarding a teddy bear, the grid diffusion model device (100) can select grid images f1, f10 and create a first mask grid image (M(1)) that places f1 in the first quadrant, f10 in the fourth quadrant, and blanks in the second and third quadrants.

[0203] Thereafter, the grid diffusion model device (100) can create a 1-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the first mask grid image (M(1)) with grid images f4 and f7 by learning the created first mask grid image (M(1)) in the 1-Step interpolation model (132).

[0204] In addition, the grid diffusion model device (100) can create a 2-Step interpolation grid image by creating n second mask grid images (where n is a natural number greater than or equal to 3) using the first image frame included in the 1-Step interpolation grid image, and then learning from the 2-Step interpolation model (133) among the interpolation models (131).

[0205] The grid diffusion model device (100) can create the n second mask grid images by considering a combination of placing each of the first image frames in the 1st and 4th quadrants and placing blanks in the 2nd and 3rd quadrants.

[0206] In addition, the grid diffusion model device (100) can create the 2-Step interpolation grid image by interpolating so that the blanks arranged in the 2nd and 3rd quadrants of the n second mask grid images are filled by the 2-Step interpolation model (133).

[0207] For the above-described 1-Step interpolation grid image having the four grid images f1, f4, f7, and f10 as the first image frame, the grid diffusion model device (100) can create a second mask grid image (M(2-1)) that places f1 in the first quadrant, f4 in the fourth quadrant, and blanks in the second and third quadrants, a second mask grid image (M(2-2)) that places f4 in the first quadrant, f7 in the fourth quadrant, and blanks in the second and third quadrants, and a second mask grid image (M(2-3)) that places f7 in the first quadrant, f10 in the fourth quadrant, and blanks in the second and third quadrants.

[0208] Thereafter, the grid diffusion model device (100) can create a 2-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the 2nd mask grid image (M(2-1)) with grid images f2 and f3 by learning the 2-Step interpolation model (133) of the 2nd mask grid image (M(2-1)).

[0209] In addition, the grid diffusion model device (100) can create a 2-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the 2nd mask grid image (M(2-2)) with grid images f5 and f6 by learning the 2-Step interpolation model (133) of the 2nd mask grid image (M(2-2)).

[0210] Finally, the grid diffusion model device (100) can create a 2-Step interpolated grid image by interpolating the blanks of the 2nd and 3rd quadrants of the 2nd mask grid image (M(2-3)) with grid images f8 and f9 by learning the 2-Step interpolation model (133) of the 2nd mask grid image (M(2-3)).

[0211] In general, the grid diffusion model device (100) can create a 2-Step interpolated grid image by interpolating between f1 and f10 with grid images f2 to f9 through learning in the 1-Step interpolation model (132) and the 2-Step interpolation model (133).

[0212] Accordingly, the grid diffusion model device (100) can output a generative video by sequentially connecting the second image frames included in the 2-Step interpolation grid image.

[0213] That is, the grid diffusion model device (100) can output a generative video of natural movement by connecting grid images f1 to f10, which are second image frames included in a 2-Step interpolation grid image.

[0214] In addition, the grid diffusion model device (100) can also continuously output the generative video after the grid image f10 by repeating the above-described process (creating the first mask grid image (M(1)), learning in the 1-Step interpolation model and the 2-Step interpolation model, etc.) by continuously selecting other internal frames of the key grid image of the teddy bear (e.g., selecting f10 and f19 and selecting f19 and f28).

[0215] At this time, the time interval between the first image frames included in the 1-Step interpolation grid image may be smaller than the time interval between the internal frames included in the key grid image, and may also be larger than the time interval between the second image frames included in the 2-Step interpolation grid image. That is, in the grid diffusion model device (100), the time interval between the second image frames included in the 2-Step interpolation grid image may be made as small as possible so that the grid images are continuously connected when outputting the generated video, thereby enabling natural video playback.

[0216] According to one embodiment of the present invention, a grid diffusion model device and method for video generation from text can be provided, which extracts a fixed number of grid images identified by text, and generates a high-quality video using these grid images using a fixed amount of GPU memory regardless of the number of frames.

[0217] Additionally, according to one embodiment of the present invention, since the dimension of a video is reduced to the dimension of an image, various image-based methods can be applied to videos, such as image manipulation and text-guided video manipulation.

[0218] Additionally, according to one embodiment of the present invention, the suitability of the model for actual video generation can be proven.

[0219] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.

[0220] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0221] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the above. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0222] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. A step of generating a key grid image corresponding to the text; and A step of creating an interpolated grid image that constitutes a video from the key grid image through an interpolation model that performs autoregressive grid image interpolation. A grid diffusion model method for video generation from text, including:

2. In paragraph 1, The steps for generating the above key grid image are: A step of receiving a prompt as the above text; A step of outputting m internal frames (where m is a natural number greater than or equal to 4) specified by the prompt from the key grid image generation model; and A step of generating the key grid image by arranging the m internal frames in chronological order. Including, The above key grid image generation model is, By selecting the inner frame of the image dimension from the video frame of the video dimension, dimension reduction is performed. A grid diffusion model method for video generation from text, including:

3. In paragraph 1, Multiple grid images (f1, … f t-1 , f t ), a step of selecting four grid images (f1 to f4) prior to the reference time point (t5) as previous grid images; A step of creating a masked grid image by blanking any two grid images among the four grid images (f5 to f8) after the above reference time point (t5) and including the remaining grid images; and A step of training the interpolation model to fill in blanks in the masked grid image based on the previous grid image. A grid diffusion model method for video generation from text, further comprising:

4. In paragraph 3, The steps for creating the above interpolation grid image are: A step of creating a first mask grid image using an internal frame included in the above key grid image, and then creating a 1-Step interpolation grid image by learning from a 1-Step interpolation model among the above interpolation models; and A step of creating a 2-Step interpolation grid image by using the first image frame included in the 1-Step interpolation grid image to create n second mask grid images (where n is a natural number greater than or equal to 3), and then learning from a 2-Step interpolation model among the interpolation models. A grid diffusion model method for video generation from text, including:

5. In paragraph 4, The steps for creating the above 1-Step interpolation grid image are: A step of creating the first mask grid image by placing internal frames included in the key grid image in quadrants 1 and 4 and blanks in quadrants 2 and 3; and A step of creating the 1-Step interpolation grid image by interpolating so that the blanks placed in the 2nd and 3rd quadrants of the first mask grid image are filled by the 1-Step interpolation model. A grid diffusion model method for video generation from text, including:

6. In paragraph 5, The steps for creating the above 2-Step interpolation grid image are: A step of creating n second mask grid images by considering a combination of placing each of the first image frames in the 1st and 4th quadrants and placing blanks in the 2nd and 3rd quadrants; and A step of creating the 2-Step interpolation grid image by interpolating so that the blanks placed in the 2nd and 3rd quadrants of the n second mask grid images are filled by the 2-Step interpolation model. A grid diffusion model method for video generation from text, including:

7. In paragraph 6, The time interval between the first image frames included in the above 1-Step interpolation grid image is A time interval smaller than the time interval between internal frames included in the above key grid image and larger than the time interval between second image frames included in the above 2-Step interpolation grid image. A grid diffusion model method for video generation from text.

8. In paragraph 4, The above grid diffusion model method, A step of outputting a generative video by sequentially connecting the second image frames included in the above 2-Step interpolation grid image. A grid diffusion model method for video generation from text, further comprising:

9. A generation unit that generates a key grid image corresponding to the text; and A processing unit that creates an interpolated grid image that constitutes a video from the key grid image through an interpolation model that performs autoregressive grid image interpolation. A grid diffusion model device for video generation from text, comprising:

10. In paragraph 9, The above generating unit, Enter the prompt as the above text, From the key grid image generation model, output m internal frames (where m is a natural number greater than or equal to 4) specified by the above prompt, The above m internal frames are arranged in chronological order to generate the above key grid image, The above key grid image generation model is, By selecting the inner frame of the image dimension from the video frame of the video dimension, dimension reduction is performed. A grid diffusion model device for video generation from text, comprising:

11. In paragraph 9, Multiple grid images (f1, … f t-1 , f t ), a model training unit that selects four grid images (f1 to f4) before a reference time point (t5) as previous grid images, blanks any two grid images among four grid images (f5 to f8) after the reference time point (t5), and creates a masked grid image including the remaining grid images, and trains the interpolation model to fill in the blanks in the masked grid image based on the previous grid images. A grid diffusion model device for generating video from text, further comprising:

12. In paragraph 11, The above processing unit, After creating a first mask grid image using the internal frame included in the above key grid image, a 1-Step interpolation grid image is created by learning from a 1-Step interpolation model among the above interpolation models. After creating n second mask grid images (where n is a natural number greater than or equal to 3) using the first image frame included in the above 1-Step interpolation grid image, a 2-Step interpolation grid image is created by learning from a 2-Step interpolation model among the above interpolation models. A grid diffusion model device for video generation from text.

13. In paragraph 12, The above processing unit, Create the first mask grid image by placing the inner frames included in the above key grid image in the first and fourth quadrants and placing the blanks in the second and third quadrants, By the above 1-Step interpolation model, the blanks placed in the 2nd and 3rd quadrants of the first mask grid image are interpolated to fill them, thereby creating the 1-Step interpolation grid image. A grid diffusion model device for video generation from text.

14. In paragraph 13, The above processing unit, Considering the combination of placing each of the first image frames in the 1st and 4th quadrants and placing the blanks in the 2nd and 3rd quadrants, the n second mask grid images are created, By the above 2-Step interpolation model, the blanks arranged in the 2nd and 3rd quadrants of the n second mask grid images are interpolated to fill them, thereby creating the 2-Step interpolation grid image. A grid diffusion model device for video generation from text.

15. In paragraph 14, The time interval between the first image frames included in the above 1-Step interpolation grid image is A time interval smaller than the time interval between internal frames included in the above key grid image and larger than the time interval between second image frames included in the above 2-Step interpolation grid image. A grid diffusion model device for video generation from text.

16. In paragraph 12, The above processing unit, A step of outputting a generative video by sequentially connecting the second image frames included in the above 2-Step interpolation grid image. A grid diffusion model device for generating video from text, further comprising:

17. A computer-readable recording medium recording a program for executing the method of paragraph 1.

Citation Information

Patent Citations

  • A data-driven automatic animation generation method and system

    CN112258608B

  • Audio video generation method and device, electronic equipment and storage medium

    CN116524898A

  • Method, system and equipment for generating movie video clip by text

    CN117478978A

  • Text video generation method and system based on potential diffusion model

    CN117729370A

  • Systems and methods for hierarchical text-conditional image generation

    US11922550B1