Method and device for generating video from text, electronic equipment and storage medium

Through parallel calculation and dynamic switching of sequence parallel dimensions, the problem of high computing resources and time consumption in Wensheng videos is solved, and the effect of efficient video generation is achieved.

CN120358398AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510838062.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

When generating video, the existing Wensheng video technology needs to alternate time, space and cross-modal attention on each frame of the screen, resulting in a large consumption of computing resources and time, affecting the generation efficiency.

Method used

The spatial attention module, the temporal attention module and the cross-modal attention module are set up in parallel, and different attention mechanisms are calculated in parallel by dynamically switching sequence parallel dimensions, controlling the broadcast frequency of different attention feature information, and reducing computing resources and time.

Benefits of technology

It significantly improves the generation efficiency of Wensheng videos, reduces computing resources and time consumption, and ensures the quality of generated videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358398A_ABST
    Figure CN120358398A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for generating a video through a text, electronic equipment and a storage medium, and relates to the technical field of computers. Due to the fact that a space attention module, a time attention module and a cross-modal attention module are arranged to calculate space attention feature information, time attention feature information and cross-modal attention feature information for a text sequence in a parallel mode, different attention mechanisms are calculated in parallel, and communication traffic can be reduced; the communication efficiency can be obviously improved; when each frame of picture is formed, the broadcast frequencies of the space attention feature information, the time attention feature information and the cross-modal attention feature information are controlled by utilizing the influence change rate of the space attention module, the time attention module and the cross-modal attention module on the generated video; according to the invention, different attention feature information can be broadcasted by adopting broadcasting mechanisms with different frequencies, so that the computing resources and time are reduced, and the video generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for generating video from text. Background Art

[0002] With the rapid development of artificial intelligence technology, Wenshengtu and Wenshengvideo tools have become important auxiliary tools for creative workers. For example, Wenshengtu Stable Diffusion model is an image generation model based on diffusion process, which can generate high-quality and high-resolution images. It gradually transforms the noise image into the target image by simulating the diffusion process. This model has strong stability and controllability, and can generate images with diversified effects and good visual effects; Wensheng video Open-Sora v has achieved a qualitative leap in the ability to generate videos from text. Users only need to enter a simple text description, and the model can convert it into a vivid and realistic video picture. This instant conversion ability from text to vision has undoubtedly injected new vitality and possibilities into the field of content creation. These tools have greatly reduced the threshold for creation and improved the efficiency of creation by converting text descriptions into visual content. However, due to the limitations of hardware equipment and video memory requirements, the time required to generate a video has increased significantly. The larger the pixels of the generated video, the higher the frame rate, and the longer it takes. Summary of the invention

[0003] The present application provides a method, device, electronic device and storage medium for generating video from text, so as to at least solve the problem in the related art that when generating video from text, time, space and cross-modal attention needs to be alternately executed for each frame, and each alternating execution needs to recalculate the attention weight, integrate local and global features in the spatial dimension, sort out the temporal dependencies of the sequence in the temporal dimension, and fuse and interact information across modalities, involving a large number of matrix operations and data processing, consuming a large amount of computing resources and time, and causing a technical problem that time consumption is greatly increased.

[0004] This application provides a method for generating a video from text, comprising: Forming a text sequence according to text information of a pre-generated video; Setting a spatial attention module, a temporal attention module and a cross-modal attention module to respectively calculate spatial attention feature information, temporal attention feature information and cross-modal attention feature information for the text sequence in parallel; Obtaining the influence change rate of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and controlling the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the influence change rate; Decode the text sequence based on the virtual twin method to generate multiple frames of images, and process each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and output a video file.

[0005] This application also provides a text-to-video device, including: A text conversion module, configured to form a text sequence according to the text information of the pre-generated video; An attention feature calculation module, configured to set a spatial attention module, a temporal attention module, and a cross-modal attention module to calculate spatial attention feature information, temporal attention feature information, and cross-modal attention feature information for the text sequence respectively in a parallel manner; A broadcast frequency control module, configured to obtain the influence change rates of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and control the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the influence change rates; An image processing module, configured to decode the text sequence based on the virtual twin method to generate multiple frames of images, and process each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and output a video file.

[0006] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above text-to-video methods when executing the computer program: Form a text sequence according to the text information of the pre-generated video; Set a spatial attention module, a temporal attention module, and a cross-modal attention module to calculate spatial attention feature information, temporal attention feature information, and cross-modal attention feature information for the text sequence respectively in a parallel manner; Obtain the influence change rates of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and control the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the influence change rates; Decode the text sequence based on the virtual twin method to generate multiple frames of images, and process each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and output a video file.

[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above text-to-video methods are implemented: Form a text sequence according to the text information of the pre-generated video; Set the spatial attention module, the temporal attention module, and the cross-modal attention module to calculate the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information for the text sequence respectively in a parallel manner; Obtain the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and control the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the change rate; Decode the text sequence to generate multiple frames of images based on the virtual twin method, and process each frame of image according to the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and output a video file.

[0008] Through the present application, since the spatial attention module, the temporal attention module, and the cross-modal attention module are set to calculate the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information for the text sequence respectively in a parallel manner, parallelly calculating different attention mechanisms can reduce the communication volume, significantly improve the communication efficiency, and when forming each frame of image, use the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video to control the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, which can implement a broadcast mechanism with different frequencies for different attention feature information to broadcast, reduce the computing resources and time, and improve the efficiency of generating the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a logic diagram of the existing DiT module; Figure 2 It is an application environment diagram of the text-to-video method in an embodiment of the present application; Figure 3Schematic flowchart of the method for generating a video from text in an embodiment of the present application; Figure 4 Logic diagram of the method for generating a video from text in an embodiment of the present application; Figure 5 Structural block diagram of the device for generating a video from text in an embodiment of the present application; Figure 6 Internal structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0011] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0012] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0013] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0014] As described in the background art, most text-to-video models adopt the Diffusion Transformer (DiT) architecture. DiT (Diffusion Transformer) is a diffusion model combined with the Transformer architecture for image and video generation tasks, which can efficiently capture dependencies in data and generate high-quality results. The core idea of DiT is to gradually add noise to the data by simulating the diffusion process and then learn to reverse this process to construct the required data samples from the noise. Taking the text-to-video model open-sora as an example, it is based on the high-quality open-source text-to-image model PixArt-α, and on this basis, a temporal attention layer is introduced to extend it to video data. Specifically, the entire architecture includes a pre-trained VAE, a text encoder, and an STDiT (SpatialTemporal Diffusion Transformer) model that utilizes spatial-temporal attention mechanisms. When predicting a video, the model randomly generates a noise, and then the noise is removed through multiple rounds of DiT, and finally a video related to the input text is obtained.

[0015] In the text-to-video model, the most time-consuming module is the DiT module, which accounts for approximately 90% of the total time consumption. Taking the opensora model as an example, the structure of each layer of its DiT is as shown in the appendix Figure 1 It stacks a one-dimensional temporal attention module on top of a two-dimensional spatial attention module in a serial manner for modeling temporal relationships. After the temporal attention module, a cross-modal attention module is used to align the semantics of the text. The DiT module of the opensora model stacks a one-dimensional temporal attention module on top of a two-dimensional spatial attention module in a serial manner. There are significant performance bottlenecks in the actual operation of this module. The main part of its time consumption is spent on 28 alternating executions of temporal, spatial, and cross-modal attention. For each alternating execution, the model needs to recalculate the attention weights, integrate local and global features in the spatial dimension, sort out the temporal dependencies of the sequence in the temporal dimension, and also perform information fusion and interaction across modalities. Such frequent alternating calculations involve a large amount of matrix operations and data processing, consuming a large amount of computing resources and time. The time consumed in this process accounts for 90% of the total time consumption, seriously affecting the running efficiency and real-time performance of the model.

[0016] The text-to-video method provided by this application can be applied to such as Figure 2In the application environment shown. Among them, the terminal 102 communicates with the server 104 through a network. The terminal 102 can input the text information of the pre-generated video into the server 104, and the server 104 processes the text information of the pre-generated video to form a video file and outputs it. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0017] As Figure 3 shown, an embodiment of the present application provides a method for generating a video from text, including the following steps: Step S1, forming a text sequence according to the text information of the pre-generated video; Step S2, setting a spatial attention module, a temporal attention module, and a cross-modal attention module to respectively calculate spatial attention feature information, temporal attention feature information, and cross-modal attention feature information for the text sequence in a parallel manner; Step S3, obtaining the influence change rate of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and controlling the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the influence change rate; Step S4, decoding the text sequence based on the virtual twin method to generate multiple frames of images, and processing each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and outputting a video file.

[0018] In this application, for the different characteristics of its three attention mechanisms: the spatial attention changes the most, involving high-frequency elements such as edges and textures; the temporal attention shows medium-frequency changes related to motion and dynamics in the video; the cross-modal attention is the most stable, connecting the text with the video content, similar to a low-frequency signal reflecting the text semantics. By adopting a broadcast mechanism with different frequencies for them, for the spatial attention mechanism, we adopt a broadcast frequency of once every two times. For the temporal attention and cross-modal attention mechanisms, because they are relatively stable, we adopt a broadcast frequency of once every four times or once every eight times. This not only improves the performance of the DiT module, but also the quality loss of the generated content can be ignored.

[0019] Through experiments, it is found that the attention differences at different time steps show a U-shaped pattern, with significant changes occurring in the initial and the last 15% of the steps, while the middle 70% of the steps are very stable with little difference.

[0020] Secondly, within the stable intermediate segment, there are differences among the attention types: Spatial attention varies the most, involving high-frequency elements such as edges and textures; Temporal attention exhibits medium-frequency variations related to motion and dynamics in the video; Cross-modal attention is the most stable, connecting text to video content, similar to a low-frequency signal reflecting text semantics.

[0021] In addition, we also use Dynamic Sequence Parallelism to compute different attention mechanisms in parallel because the computations of each attention mechanism's sequence dimension are independent of each other. For example, when calculating the spatial transformer block, it has nothing to do with calculating the temporal sequence dimension, and the data can be split among machines without affecting the results. To utilize this feature, we propose to dynamically switch the dimensions of sequence parallelism. In this way, there is no need for complex communication within the attention module, and only one dynamic switch (Dynamic Switch) is required between different computing stages, greatly reducing the communication overhead. Relying on this multi-dimensional computing feature, our method can be extended to all multi-dimensional transformers.

[0022] In this embodiment, obtaining the change rates of the spatial attention module, temporal attention module, and cross-modal attention module on the generated video, and controlling the broadcast frequencies of the spatial attention feature information, temporal attention feature information, and cross-modal attention feature information in descending order of the change rates includes: Analyzing the change rates of the spatial attention module, temporal attention module, and cross-modal attention module on the generated video; Setting the influence weights of the spatial attention module, temporal attention module, and cross-modal attention module according to the change rates; Setting the broadcast frequencies of the spatial attention feature information, temporal attention feature information, and cross-modal attention feature information at each time step according to the influence weights of the spatial attention module, temporal attention module, and cross-modal attention module respectively.

[0023] In this embodiment, analyzing the change rates of the spatial attention module, temporal attention module, and cross-modal attention module on the generated video; setting the influence weights of the spatial attention module, temporal attention module, and cross-modal attention module according to the change rates includes: Obtaining the amount of change in the frame information between two adjacent frames when the spatial attention module, temporal attention module, and cross-modal attention module generate the video, and calculating the change rate of the influence on the generated video based on the amount of change in the frame information between two adjacent frames; Respectively counting the number of consecutive frames when the change rates corresponding to the spatial attention module, temporal attention module, and cross-modal attention module are greater than the first threshold; Set the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the number of consecutive frames when the influence change rate is greater than the first threshold. The value of the influence weight is positively correlated with the number of consecutive frames when the influence change rate is greater than the first threshold.

[0024] Among them, a broadcast controller can be added, and this broadcast controller is responsible for controlling the broadcast frequency of each attention module.

[0025] It can be understood that setting the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information also includes: Identify the video content complexity according to the text information of the pre-generated video; Adjust the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the video content complexity; During the process of forming the video, obtain the computing resource load rate; Adjust the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the computing resource load rate.

[0026] Among them, the algorithm for dynamically adjusting the broadcast frequency according to the video content complexity and the model load makes the acceleration strategy more intelligent and flexible. For example, when processing a video of a complex scene, automatically increase the broadcast frequency of the cross-modal attention module; while in a simple scene, reduce the frequency, which can further improve the efficiency while ensuring the quality.

[0027] In this embodiment, decode the text sequence based on the virtual twin method to generate multiple frames of images, and process each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information. The output video file includes: Obtain the sequence of multiple frames of images for decoding the text sequence to generate a video, and each frame of image generated corresponds to a time step; Obtain the value of the broadcast frequency of the spatial attention feature information, obtain the value of the broadcast frequency of the temporal attention feature information, obtain the value of the broadcast frequency of the cross-modal attention feature information, and form a broadcast frequency control matrix according to the values of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information; At each time step, control the time steps for the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information to be broadcast and propagated according to the broadcast frequency control matrix; Process the image at each time step according to the broadcast spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and form and output a video file in the order of the sequence of multiple frames of images for generating the video.

[0028] For the video generated by the video generation model optimized by the present invention, the FID error is calculated frame by frame with the video generated before optimization, and the FID is about 0.005. This indicates that while achieving acceleration, the quality of the generated video is ensured, which is an important advantage and a key consideration of this solution.

[0029] In this embodiment, at each time step, the time steps for controlling the broadcast propagation of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the broadcast frequency control matrix include: Let the current time step be X t , and obtain a multi-frame picture sequence according to the time sequence of multiple frames of pictures of the generated video; Obtain the broadcast frequency control matrix [A, B, C], where A is the value of the broadcast frequency of the spatial attention feature information, B is the value of the broadcast frequency of the temporal attention feature information, and C is the value of the broadcast frequency of the cross-modal attention feature information; In response to the time sequence of multiple frames of pictures of the generated video being in reverse order, set the time steps for starting the broadcast of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information as X t-nA , X t-nB , X t-nC , where n is an integer; In response to the time sequence of multiple frames of pictures of the generated video being in positive order, set the time steps for starting the broadcast of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information as X t+nA , X t+nB , X t+nC , where n is an integer; According to the time sequence of multiple frames of pictures of the generated video, control the spatial attention feature information to be sequentially propagated from the current time step X t to the picture corresponding to the subsequent A - 1 time steps, control the temporal attention feature information to be sequentially propagated from the current time step X t to the picture corresponding to the subsequent B - 1 time steps, and control the cross-modal attention feature information to be sequentially propagated from the current time step X t to the picture corresponding to the subsequent C - 1 time steps.

[0030] Specifically, according to the time sequence of multiple frames of pictures of the generated video, control the spatial attention feature information to be sequentially propagated from the current time step X t to the picture corresponding to the subsequent A - 1 time steps, control the temporal attention feature information to be sequentially propagated from the current time step X tPropagate sequentially to the frames corresponding to the subsequent B - 1 time steps, and control the cross - modal attention feature information to be respectively from the current time step X t Propagating sequentially to the frames corresponding to the subsequent C - 1 time steps includes: In response to the time order of the multiple frames of the generated video being in reverse order, control the spatial attention feature information, temporal attention feature information, and cross - modal attention feature information to be respectively from the current time step X t Propagate sequentially to time step X t-A-1 、X t-B-1 、X t-C-1 ; In response to the time order of the multiple frames of the generated video being in forward order, control the spatial attention feature information, temporal attention feature information, and cross - modal attention feature information to be respectively from the current time step X t Propagate sequentially to time step X t+A-1 、X t+B-1 、X t+C-1 .

[0031] Among them, attention broadcasting is used to reduce unnecessary attention calculations. In the middle part, the attention shows slight differences. The study broadcasts the attention output of one diffusion step to several subsequent steps, thus significantly reducing the computational cost, as shown in the appendix Figure 4 . Among them, Dynamic Switch only needs one All2All operation to complete the switching of parallel dimensions, ensuring the normal execution of subsequent calculations. By such a method, compared with the previous sequential parallel work, the communication volume can be reduced by more than 75%, and the communication efficiency is very significantly improved.

[0032] When the broadcast controller manages the broadcast frequency control matrix [A, B, C], the pseudo - code of the broadcast controller is as follows: class DiTBroadcastManager(Config): def __init__( self, spatial_broadcast: bool = True, spatial_threshold: list = [15,31], spatial_range: int = 2,# For spatial_attn, its broadcast frequency is 2 temporal_broadcast: bool = True, temporal_threshold: list = [15,31], temporal_range: int = 4, #For temporal_attn, the broadcast frequency is 4 cross_broadcast : bool = True, cross_threshold: list = [15,31], cross_range: int = 4, #For cross_attn, the broadcast frequency is 4 mlp_broadcast : bool = False ).

[0033] like Figure 4 As shown, the broadcast frequency control matrix [A, B, C] = [2, 4, 4], the function implemented by the above pseudo code is: Assume that the current time step is X t , the time order of the multi-frame images of the generated video is in reverse order. When the spatial attention module adopts a broadcast mechanism with a broadcast frequency of 2, then X t-1 It is by X t For the temporal attention module and the cross-modal attention module, a broadcasting mechanism with a broadcasting frequency of 4 is adopted, then X t-1 , X t-2 , X t-3 By X t The broadcast gets, the smaller the change in attention, the greater the broadcast frequency.

[0034] In this embodiment, processing the picture at each time step according to the broadcasted spatial attention feature information, temporal attention feature information, and cross-modal attention feature information includes: Determine whether there is broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step; In response to the presence of broadcasted spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step, obtaining the broadcasted spatial attention feature information, temporal attention feature information, or cross-modal attention feature information; In response to the absence of broadcast spatial attention feature information, temporal attention feature information or cross-modal attention feature information at the current time step, the spatial attention feature information, temporal attention feature information or cross-modal attention feature information is recalculated by the spatial attention module, the temporal attention module and the cross-modal attention module, and the recalculated spatial attention feature information, temporal attention feature information or cross-modal attention feature information is used in the picture at the current time step and broadcast.

[0035] Specifically, in response to the existence of broadcast spatial attention feature information at the current time step, the value of the spatial attention feature information at the previous time step is taken, otherwise the spatial attention feature information is recalculated and the recalculated spatial attention feature information is used in the picture at the current time step and broadcasted; In response to the existence of time attention feature information to be broadcasted at the current time step, the value of the time attention feature information of the previous time step is taken, otherwise the time attention feature information is recalculated and the recalculated time attention feature information is used in the picture of the current time step and broadcasted; In response to the existence of cross-modal attention feature information to be broadcast at the current time step, the value of the cross-modal attention feature information of the previous time step is taken, otherwise the cross-modal attention feature information is recalculated and the recalculated cross-modal attention feature information is used in the picture of the current time step and broadcast.

[0036] In this embodiment, when broadcasting the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, it also includes: Obtain the spatial attention feature information, temporal attention feature information, and cross-modal attention feature information broadcast at the current time step; The spatial attention feature information, temporal attention feature information, and cross-modal attention feature information to be broadcasted are set to a grouped broadcast mode or a layered broadcast mode.

[0037] Among them, in addition to the single broadcast mechanism, attempts are made to combine different types of broadcast mechanisms, such as packet broadcast mode and layered broadcast mode, to tap greater acceleration potential.

[0038] The actual application pseudo code is as follows: def forward( self, x, y, t, mask=None,# text mask x_mask=None, # temporal mask t0=None,# t with timestamp=0 T=None,# number of frames S=None,# number of pixel patches timestep=None, all_timesteps=None, ): B, N, C = x.shape shift_msa, scale_msa, gate_msa, shift_mlp, scale_mlp, gate_mlp = ( self.scale_shift_table[None] + t.reshape(B, 6, -1) ).chunk(6, dim=1) if x_mask is not None: shift_msa_zero, scale_msa_zero, gate_msa_zero, shift_mlp_zero, scale_mlp_zero, gate_mlp_zero = ( self.scale_shift_table[None] + t0.reshape(B, 6, -1) ).chunk(6, dim=1) if enable_broadcast(): if self.temporal: broadcast_attn, self.attn_count = if_broadcast_temporal(int(timestep[0]), self.attn_count) else: broadcast_attn, self.attn_count = if_broadcast_spatial(int(timestep[0]), self.attn_count) if enable_broadcast() and broadcast_attn: x_m_s = self.last_attn else: x_m = t2i_modulate(self.norm1(x), shift_msa, scale_msa)#torch.Size([2, 50400, 1152]) if x_mask is not None: x_m_zero = t2i_modulate(self.norm1(x), shift_msa_zero, scale_msa_zero) x_m = self.t_mask_select(x_mask, x_m, x_m_zero, T, S) # attention if self.temporal: if self.parallel_manager.sp_size>1: x_m, S, T = self.dynamic_switch(x_m, S, T, to_spatial_shard=True) x_m = rearrange(x_m, "B (T S) C ->(B S) T C", T=T, S=S) x_m = self.attn(x_m) x_m = rearrange(x_m, "(B S) T C ->B (T S) C", T=T, S=S) if self.parallel_manager.sp_size>1: x_m, S, T = self.dynamic_switch(x_m, S, T, to_spatial_shard=False) else: x_m = rearrange(x_m, "B (T S) C ->(B T) S C", T=T, S=S) x_m = self.attn(x_m) x_m = rearrange(x_m, "(B T) S C ->B (T S) C", T=T, S=S) x_m_s = gate_msa * x_m if x_mask is not None: x_m_s_zero = gate_msa_zero * x_m x_m_s = self.t_mask_select(x_mask, x_m_s, x_m_s_zero, T, S) if enable_broadcast(): self.last_attn = x_m_s # residual x = x + self.drop_path(x_m_s) # cross attention if enable_broadcast(): broadcast_cross, self.cross_count = if_broadcast_cross(int(timestep[0]), self.cross_count) if enable_broadcast() and broadcast_cross: x = x + self.last_cross else: x_cross = self.cross_attn(x, y, mask) if enable_broadcast(): self.last_cross = x_cross x = x + x_cross.

[0039] The main function of the above code is to determine whether enable_broadcast is True at the current time step. A judgment is made before the calculation of the spatial attention module, the temporal attention module, and the cross-modal attention module. If it is True, the value of the previous time step is directly taken. If it is False, an attention calculation is performed and the attention feature information to be broadcast is updated.

[0040] The frequency controller module can be modified later to modify the broadcast frequency of temporal attention, spatial attention, and cross-modal attention in the video generation model. For different video models, the broadcast frequency of the attention mechanism can be fine-tuned to achieve the best results in terms of video generation efficiency and quality.

[0041] In addition, the present invention also uses DSP technology to improve the communication performance between different attention modules. DSP is a technology for efficient sequence parallel computing to implement a multi-dimensional Transformer model. The key idea is to dynamically switch the dimension of parallel computing according to the current computing stage, taking advantage of the potential characteristics of multi-dimensional attention. This dynamic dimension switching allows for sequence parallel computing with minimal communication overhead compared to applying traditional single-dimensional parallel computing to multi-dimensional models.

[0042] In this solution, the spatial attention module, the temporal attention module, and the cross-modal attention module are three independent modules that can be applied to DSP technology. Through DSP technology, these three modules can be executed in parallel. Tests have found that the performance of our video generation can be improved by about 40% after using DSP technology.

[0043] In this embodiment, forming a text sequence according to the text information of the pre-generated video includes: Obtaining the text information of the pre-generated video and encoding the text information to form a text encoding; Denosing the noise of the text encoding to form a text sequence.

[0044] In the above text-to-video method, since the spatial attention module, the temporal attention module, and the cross-modal attention module are set to calculate the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information of the text sequence in parallel respectively, parallel computing of different attention mechanisms can reduce the communication volume, significantly improve the communication efficiency, and when forming each frame of the picture, use the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video to control the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, so as to implement a broadcast mechanism with different frequencies for different attention feature information, reduce computing resources and time, and improve the efficiency of generating the video.

[0045] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0046] In one embodiment, as Figure 5 shown, a text-to-video device 10 is provided, including: a text conversion module 1, an attention feature calculation module 2, a broadcast frequency control module 3, and a picture processing module 4.

[0047] The text conversion module 1 is used to form a text sequence according to the text information of the pre-generated video.

[0048] The attention feature calculation module 2 is used to set the spatial attention module, the temporal attention module, and the cross-modal attention module to calculate the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information for the text sequence in parallel.

[0049] The broadcast frequency control module 3 is used to obtain the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and control the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the change rate of influence.

[0050] The picture processing module 4 is used to decode the text sequence based on the virtual twin method to generate multiple frames of pictures, and process each frame of picture according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and output a video file.

[0051] In this embodiment, obtaining the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and controlling the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information in a descending order of the change rate of influence includes: Analyzing the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video; Setting the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the change rate of influence; Setting the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information at each time step respectively according to the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module.

[0052] In this embodiment, analyzing the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video; setting the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the change rate of influence includes: Obtaining the amount of change in the picture information of adjacent two frames of pictures when the spatial attention module, the temporal attention module, and the cross-modal attention module generate the video, and calculating the change rate of the influence on the generated video according to the amount of change in the picture information of adjacent two frames of pictures; Respectively counting the number of consecutive frames of pictures when the influence change rates corresponding to the spatial attention module, the temporal attention module, and the cross-modal attention module are greater than the first threshold; Set the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the number of consecutive frames when the influence change rate is greater than the first threshold. The value of the influence weight is positively correlated with the number of consecutive frames when the influence change rate is greater than the first threshold.

[0053] In this embodiment, based on the virtual twin method, the text sequence is decoded to generate multiple frames of images, and each frame of the image is processed according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information. The output video file includes: Obtain a sequence of multiple frames of images generated by decoding the text sequence into a video. Each frame of the generated image corresponds to a time step. Obtain the value of the broadcast frequency of the spatial attention feature information, obtain the value of the broadcast frequency of the temporal attention feature information, obtain the value of the broadcast frequency of the cross-modal attention feature information, and form a broadcast frequency control matrix according to the values of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information. At each time step, control the time steps for the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information to be broadcast and propagated according to the broadcast frequency control matrix. Process the image at each time step according to the broadcast spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and form and output a video file in the order of the sequence of multiple frames of images generated by the video.

[0054] In this embodiment, at each time step, controlling the time steps for the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information to be broadcast and propagated according to the broadcast frequency control matrix includes: Let the current time step be X t , and obtain a sequence of multiple frames of images according to the time order of the multiple frames of images generated by the video. Obtain the broadcast frequency control matrix [A, B, C], where A is the value of the broadcast frequency of the spatial attention feature information, B is the value of the broadcast frequency of the temporal attention feature information, and C is the value of the broadcast frequency of the cross-modal attention feature information. In response to the time order of the multiple frames of images generated by the video being in reverse order, set the time steps for the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information to start broadcasting as X t-nA , X t-nB , X t-nC , where n is an integer. In response to the time sequence of multiple frames of the generated video being in ascending order, set the time steps for starting the broadcast of the spatial attention feature information, temporal attention feature information, and cross-modal attention feature information as X according to the broadcast frequency control matrix [A, B, C]. t+nA 、X t+nB 、X t+nC , where n is an integer; According to the time sequence of multiple frames of the generated video, control the spatial attention feature information to sequentially propagate from the current time step X t to the corresponding frame at the subsequent A - 1 time steps in sequence, control the temporal attention feature information to sequentially propagate from the current time step X t to the corresponding frame at the subsequent B - 1 time steps in sequence, and control the cross-modal attention feature information to sequentially propagate from the current time step X t to the corresponding frame at the subsequent C - 1 time steps in sequence.

[0055] In this embodiment, processing the frame at each time step according to the broadcast spatial attention feature information, temporal attention feature information, and cross-modal attention feature information includes: Judge whether there is broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step; In response to the existence of broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step, obtain the broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information; In response to the non-existence of broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step, recalculate the spatial attention feature information, temporal attention feature information, or cross-modal attention feature information through the spatial attention module, temporal attention module, and cross-modal attention module, and use the recalculated spatial attention feature information, temporal attention feature information, or cross-modal attention feature information for the frame at the current time step and broadcast it.

[0056] In this embodiment, forming a text sequence according to the text information of the pre-generated video includes: Obtain the text information of the pre-generated video, and encode the text information to form a text encoding; Denoise the noise of the text encoding to form a text sequence.

[0057] In the above text-to-video generation device, since the spatial attention module, the temporal attention module, and the cross-modal attention module calculate the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information for the text sequence in parallel, respectively, calculating different attention mechanisms in parallel can reduce the communication volume, significantly improve the communication efficiency, and when forming each frame of the video, use the spatial attention module, the temporal attention module, and the cross-modal attention module to control the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the rate of change of the influence on the generated video, so as to implement a broadcast mechanism with different frequencies for different attention feature information, reducing the computing resources and time, and improving the efficiency of generating the video.

[0058] For the description of the features in the corresponding embodiments of the text-to-video generation device, reference can be made to the relevant descriptions in the corresponding embodiments of the text-to-video generation method, which will not be elaborated here one by one.

[0059] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above text-to-video generation method embodiments.

[0060] In one embodiment, the electronic device may be a server, and its internal structure diagram may be as Figure 6 shown. The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store text-to-video data. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a text-to-video generation method.

[0061] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above text-to-video generation method embodiments when running.

[0062] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other media that can store computer programs.

[0063] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the text-to-video method are implemented.

[0064] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the text-to-video method are implemented.

[0065] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0066] The above has introduced in detail a text-to-video method, device, electronic device, and storage medium provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A method for generating a video from text, characterized in that, Including: Form a text sequence according to the text information of the pre-generated video; Set up a spatial attention module, a temporal attention module, and a cross-modal attention module to calculate spatial attention feature information, temporal attention feature information, and cross-modal attention feature information for the text sequence respectively in a parallel manner; Obtain the influence change rate of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and control the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the influence change rate; Decode the text sequence to generate multiple frames of images based on the virtual twin method, and process each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and output a video file.

2. The method for generating a video according to claim 1, wherein The obtaining the influence change rate of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and controlling the broadcast frequency of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the influence change rate includes: Analyze the influence change rate of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video; Set the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the influence change rate; Set the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information at each time step respectively according to the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module.

3. The method for generating a video according to claim 2, wherein The analyzing the influence change rate of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video; The setting the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the influence change rate includes: Obtain the amount of change in the frame information between two adjacent frames when the spatial attention module, the temporal attention module, and the cross-modal attention module generate the video, and calculate the influence change rate on the generated video according to the amount of change in the frame information between the two adjacent frames; Respectively count the number of consecutive frames of images when the influence change rates corresponding to the spatial attention module, the temporal attention module, and the cross-modal attention module are greater than a first threshold; Set the influence weights of the spatial attention module, the temporal attention module, and the cross-modal attention module according to the number of consecutive frames of images when the influence change rate is greater than the first threshold, and the value of the influence weight is positively correlated with the number of consecutive frames of images when the influence change rate is greater than the first threshold.

4. The method for generating a video according to claim 1, wherein The method based on the virtual twin method decodes the text sequence to generate multiple frames of images, and processes each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information. The output video file includes: Obtain a sequence of multiple frames of images for decoding the text sequence into a video, and assign a time step to each generated frame of image; Obtain the value of the broadcast frequency of the spatial attention feature information, obtain the value of the broadcast frequency of the temporal attention feature information, obtain the value of the broadcast frequency of the cross-modal attention feature information, and form a broadcast frequency control matrix according to the values of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information; At each time step, control the time steps for the broadcast propagation of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the broadcast frequency control matrix; Process the images at each time step according to the broadcast spatial attention feature information, temporal attention feature information, and cross-modal attention feature information, and form and output a video file in the order of the sequence of multiple frames of images of the generated video.

5. The method for generating a video according to claim 4, wherein The step of, at each time step, controlling the time steps for the broadcast propagation of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the broadcast frequency control matrix includes: Let the current time step be X t , and obtain a multi-frame sequence according to the time order of multiple frames of the generated video Obtain the broadcast frequency control matrix [A, B, C], where A is the value of the broadcast frequency of the spatial attention feature information, B is the value of the broadcast frequency of the temporal attention feature information, and C is the value of the broadcast frequency of the cross-modal attention feature information; In response to the time sequence of multiple frames of the generated video being in reverse order, the time steps at which the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information are respectively set to start broadcasting according to the broadcast frequency control matrix [A, B, C] are X t-nA , X t-nB , X t-nC , where n is an integer; In response to the time sequence of generating multiple frames of the video being in the positive order, set the time steps for starting the broadcast of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the broadcast frequency control matrix [A, B, C] to be X t+nA , X t+nB , X t+nC , where n is an integer; According to the time sequence of multiple frames of the generated video, control the spatial attention feature information from the current time step X t and sequentially propagate it to the frames corresponding to the subsequent A - 1 time steps, control the temporal attention feature information from the current time step X t and sequentially propagate it to the frames corresponding to the subsequent B - 1 time steps, control the cross-modal attention feature information respectively from the current time step X t and sequentially propagate it to the frames corresponding to the subsequent C - 1 time steps.

6. The method for generating a video according to claim 5, wherein The step of processing the images at each time step according to the broadcast spatial attention feature information, temporal attention feature information, and cross-modal attention feature information includes: Determine whether there is broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step; In response to the existence of broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step, obtain the broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information; In response to the non-existence of broadcast spatial attention feature information, temporal attention feature information, or cross-modal attention feature information at the current time step, recalculate the spatial attention feature information, temporal attention feature information, or cross-modal attention feature information through the spatial attention module, the temporal attention module, and the cross-modal attention module, and use the recalculated spatial attention feature information, temporal attention feature information, or cross-modal attention feature information in the image at the current time step and perform broadcasting.

7. The method for generating a video according to the text described in claim 1, wherein The step of forming a text sequence according to the text information of the pre-generated video includes: Obtain the text information of the pre-generated video, and encode the text information to form a text encoding; Denoise the noise of the text encoding to form a text sequence.

8. A text generation video device, characterized in that, Includes: A text conversion module for forming a text sequence according to the text information of a pre-generated video; An attention feature calculation module for setting a spatial attention module, a temporal attention module, and a cross-modal attention module to calculate spatial attention feature information, temporal attention feature information, and cross-modal attention feature information for the text sequence respectively in a parallel manner; A broadcast frequency control module for obtaining the change rate of the influence of the spatial attention module, the temporal attention module, and the cross-modal attention module on the generated video, and controlling the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information according to the change rate; A frame processing module for decoding the text sequence to generate multiple frames of images based on the virtual twin method, and processing each frame of image according to the broadcast frequencies of the spatial attention feature information, the temporal attention feature information, and the cross-modal attention feature information, and outputting a video file.

9. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the text-to-video method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the text-to-video method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video generation model training method and device, equipment and storage medium

    CN117499711A

  • Multi-frequency analysis video generation method based on attention mechanism

    CN117611468A

  • Multi-modal data fusion control method and device, equipment and medium

    CN118734250A

  • AIGC multi-mode audio-visual content creation method and system

    CN119299806A

  • Intention identification method, apparatus and device based on attention mechanism, and storage medium

    WO2021232589A1