Video generation method, device, storage medium and program product

Through the regressive long video generation system, the feature injector and multi-module work together, the spatial content error and timing discontinuity of long video generation in the Wensheng video model are solved, and high-quality long video generation is achieved under efficient and low computing power.

CN120416623BActive Publication Date: 2025-08-26SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510874232.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-08-26
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing Wensheng video model is prone to problems of spatial content errors and timing discontinuity when generating long videos, especially as time increases, it is difficult to maintain the forgetting and discontinuity of video frames.

Method used

A regressive long video generation system is adopted, and the feature injector is used to output the characteristics of the target video frame, combined with the historical frame condition attention module, long-term memory perception module and randomly enhanced hybrid inference module, high-quality long videos are generated through parallel hierarchical generation.

Benefits of technology

Maintain the continuity and consistency of video frames, prevent forgetting, and realize the generation of high-quality long videos, reduce computing power consumption, and improve generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416623B_ABST
    Figure CN120416623B_ABST
Patent Text Reader

Abstract

The present application discloses a video generation method, device, storage medium, and program product, which relate to the field of artificial intelligence technology, including: screening target video frames from the current first video and outputting a first feature vector using a feature injector; using a Vincent video model to generate a second feature vector based on preset noise and the description text of the current first video to generate corresponding conditional features; determining the anchor frame of the current first video, generating a target vector based on the anchor frame and the description text, and further generating a fusion feature; generating a new current first video based on the fusion feature, jumping to the step of screening target video frames, until the number of frames of each first video meets the preset conditions, and generating the corresponding second video. By capturing the visual features of historical frames and determining the fusion features based on the anchor frame and the description text, forgetting can be prevented during the generation of future video frames, solving problems such as feature forgetting and temporal discontinuity, and maintaining the consistency and smoothness of the generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a video generation method, device, storage medium, and program product. Background Art

[0002] The AIGC (Artificial Intelligence Generated Content) movement is rapidly gaining momentum, and OpenAI recently released SoRA, a large-scale model for video generation. The SoRA model can generate high-definition videos from text descriptions, offering a new approach for rapid prototyping and proof-of-concept in creative content production, film, animation, gaming, and advertising.

[0003] While current video models can generate short videos that follow textual content, they primarily use a cascaded approach to continuously generate new videos. Consequently, because the currently generated video frame is used as input to predict future frames, this often leads to forgetting over time. Generating new video frames not only results in spatial content errors (e.g., changes in the dog's facial features and color in the new video frame), but also in discontinuities and dramatic jumps in timing. Therefore, generating high-quality long videos is a worthy research topic. Summary of the Invention

[0004] The present application provides a video generation method, device, storage medium and program product to at least solve the problem in the related art that the prediction of future frames of a video generates content errors and temporal discontinuities as time increases.

[0005] This application provides a video generation method, comprising:

[0006] Filtering a preset number of target video frames from the current first video, and outputting a plurality of first feature vectors corresponding to the target video frames using a preset feature injector;

[0007] Generate a plurality of second feature vectors based on the preset noise and the description text corresponding to the current first video using a preset Wensheng video model, and generate corresponding conditional features based on the first feature vector and the second feature vector;

[0008] Determine an anchor frame in the current first video, generate a corresponding target vector according to the anchor frame and the description text, and determine a corresponding fusion feature based on the target vector, the first feature vector, and the conditional feature;

[0009] A new current first video is generated based on the fusion features, and the process jumps to the step of filtering out a preset number of target video frames from the current first video until the sum of the number of frames of each first video meets the preset frame number condition, and then a corresponding second video is generated based on each first video.

[0010] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned video generation methods when executing the computer program.

[0011] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned video generation methods are implemented.

[0012] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned video generation methods when executed by a processor.

[0013] Through this application, the feature injector can be used to output the features of the target video frame, thereby capturing the visual features of the historical frame and maintaining the continuity of the future generated video frames. The anchor frame in the current video can be determined, and the corresponding target vector is generated according to the anchor frame and the description text, and the fusion feature is determined, so that high-level scene semantics can be captured from the earlier history and forgetting can be prevented in the process of future video frame generation. Since the currently generated video frame is used as the next input to predict the future frame of the video, it is possible to solve problems such as feature forgetting and temporal discontinuity, and maintain the consistency and smoothness of the generated long video. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 A flow chart of a video generation method provided in an embodiment of the present application;

[0016] Figure 2 A regression-based long video generation network framework diagram provided in an embodiment of the present application;

[0017] Figure 3 A regression-based long video generation network framework diagram provided in an embodiment of the present application;

[0018] Figure 4 A network structure diagram of an attention model provided in an embodiment of the present application;

[0019] Figure 5 A schematic diagram of splitting a long video provided in an embodiment of the present application;

[0020] Figure 6 A network diagram of a key frame generation diffusion model provided in an embodiment of the present application;

[0021] Figure 7 A schematic diagram of a cross-attention sampling layer provided in an embodiment of the present application;

[0022] Figure 8 A network diagram of a key frame interpolation diffusion model provided in an embodiment of the present application;

[0023] Figure 9 A flowchart for inferring and generating short video content provided in an embodiment of the present application;

[0024] Figure 10 A schematic diagram of a key frame extraction network model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0027] The current Wensheng video model mainly uses a method that can cascade short videos to continuously generate new videos. Specifically, the currently generated video frame is used as the next input to predict the future frame of the video. However, as time goes by, spatial content errors and temporal discontinuities usually occur. The present application can use the feature injector to output the features of the target video frame, thereby capturing the visual features of the historical frame and maintaining the continuity of the future generated video frames. It can also determine the anchor frame in the current first video, generate the corresponding target vector based on the anchor frame and the description text, and then determine the fusion feature, thereby preventing forgetting in the future video frame generation process and maintaining the consistency and smoothness of the generated long video.

[0028] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] It should be pointed out that this application is mainly used in AIGC's Vincent video scenarios. Among them, it can be understood that the input of the Vincent video task is a text description, and the output is a video that matches the text description content. The current mainstream Vincent video model is limited by computing power and can usually only produce short videos of less than 30 frames. This application proposes a fast long video content generation system architecture, which only needs to use the traditional Vincent graph model and adopt a regression-based generation method to generate a long video generation model (more than 1000 frames). Through the regression-based generation method of the generated content system architecture in this application, a parallel hierarchical parallel generation method can be adopted, which greatly improves the generation efficiency.

[0030] Specifically, an embodiment of the present application provides a video generation method, including:

[0031] Step S11: Filter out a preset number of target video frames from the current first video, and use a preset feature injector to output a plurality of first feature vectors corresponding to the target video frames.

[0032] In this embodiment, Figure 2 As shown, a preset number of target video frames need to be screened out from the current first video. Specifically, based on a preset descriptive text, for example, "a person is skating on an ice-covered island full of trees", the corresponding video clip, i.e., the current first video, can be generated using a pre-trained keyframe vignette model and a keyframe interpolation model. For example, a video clip that is consistent with the input text is inferred using a keyframe vignette model and a keyframe interpolation model. However, it should be noted that the video clip at this time is usually only about 2 seconds long. Figure 2 The regressive long video generation network framework shown can regressively generate long video content (videos of 1 minute or even longer) starting from the initial content video frames (32 frames). It should be pointed out that this embodiment adds a historical frame conditional attention module, a long-term memory perception module, and a random enhanced hybrid reasoning module to the traditional Vincent video model (such as the ModelScope model). Among them, the network modules that need to be trained are the historical frame conditional attention module, the MLP (Multilayer Perceptron) module in the long-term memory perception module, and the convolution module. Therefore, the training parameters account for a very small proportion compared to the Vincent video model, resulting in extremely low computing power consumption, but can generate high-quality long video content.

[0033] Among them, the historical frame conditional attention module injects the tail content of the previously generated video content into the Vincent video model as a condition, so that it can generate videos with consistent temporal content when generating new video content; the long-term memory perception module ensures the consistency of spatial content in the process of generating new video content, without losing the details of objects in the video; the random enhanced hybrid inference module does not require training any network parameters, and can obtain high-quality long videos through inference through a hybrid sampling strategy combined with a high-quality Vincent video model.

[0034] Specifically, after obtaining the current first video content V (video length 32 frames) through the keyframe viz-graph model and the keyframe interpolation model, a preset feature injector is used to output several first feature vectors corresponding to the target video frame. Specifically, the target video frame can be processed using a preset image encoder and convolutional layers, and Gaussian noise is added to the processed target video frame to obtain a noisy target video frame. The noisy target video frame is then input into the preset feature injector to output several levels of first feature vectors corresponding to the target video frame. These levels are divided based on the network structure of the preset image encoder, and the network structure and initialization parameters of the preset image encoder are the same as those of the preset feature injector.

[0035] like Figure 2 As shown, the current first video in this embodiment can be the video content output by the key frame interpolation model, or it can be the video content generated by the long video generation framework in this embodiment, and then the last 16 frames of the video content are captured. , and the initial video condition vectors of each stage are obtained through the historical frame condition attention module. It should be pointed out that, Figure 2 As shown in Figure 1, the historical frame conditional attention module in this embodiment consists of an image encoder and a feature injector. The image encoder uses the ViT-L (Vision Transformer Low resolution, a low-resolution variant of the ViT model) network structure (not limited to a specific type).

[0036] Step S12: using a preset Wensheng video model to generate a plurality of second feature vectors based on preset noise and the description text corresponding to the current first video, and generating corresponding conditional features based on the first feature vector and the second feature vector.

[0037] In this embodiment, a preset Wensheng video model can be used to generate a plurality of second feature vectors based on the preset noise and the description text corresponding to the current first video. , and based on the first eigenvector and the second eigenvector Generate corresponding conditional features Specifically, After the image encoder and zero convolution (the convolution layer is randomly initialized with a parameter mean of 0) and Gaussian noise is added, and then the feature vectors of each stage are obtained through the feature injector. (i represents the index of the network stage, such as Figure 2 As shown in the figure, there are 4 stages). It can be understood that the feature injector network structure and Figure 2 The encoder of the Chinese video model is the same, and the initialization parameters also load the network weights of the encoder of the Chinese video model. The feature vectors of each stage obtained by the Vincent video model encoder Input to the attention model.

[0038] That is to say, when generating corresponding conditional features based on the first eigenvector and the second eigenvector in this embodiment, the second eigenvector can be normalized using the preset spatiotemporal normalization layer, and the normalized second eigenvector can be processed using the first preset fully connected layer, and then the first eigenvector and the processed second eigenvector are fused through the preset temporal attention layer, and the fused eigenvector is processed using the second preset fully connected layer to obtain the conditional features. Specifically, the network structure of the above attention model is as follows: Figure 3 As shown, After the spatiotemporal normalization layer Norm and the fully connected layer ,and Pass through the temporal attention layer together , then through the fully connected layer It is worth noting that the above temporal attention layer includes 、 、 The parameters of the three fully connected layers are as follows: is the dimension of the key vector, used to scale the dot product:

[0039] ;

[0040] Accordingly, Figure 3 The specific formula of the attention model network structure diagram shown is as follows:

[0041] ;

[0042] ;

[0043] ;

[0044] .

[0045] Step S13: determine the anchor frame in the current first video, generate a corresponding target vector according to the anchor frame and the description text, and determine the corresponding fusion feature based on the target vector, the first feature vector and the conditional feature.

[0046] In this embodiment, the anchor frame in the current first video can be first determined, a corresponding target vector can be generated based on the anchor frame and the description text, and a corresponding fusion feature can be determined based on the target vector, the first feature vector and the conditional feature. Specifically, when generating the corresponding target vector based on the anchor frame and the description text, the description text can be encoded using the text encoder of the preset multimodal training model to obtain the corresponding text feature vector, and the anchor frame can be encoded using the image encoder of the preset multimodal training model to obtain the corresponding image feature vector, and then the image feature vector can be feature mapped using the preset multi-layer perceptron according to the vector dimension of the text feature vector, and the text feature vector and the mapped image feature vector are spliced ​​to obtain a spliced ​​feature vector, and then the weights corresponding to each level of the preset Wensheng video model are determined, and the corresponding target vector is generated based on the weights and the spliced ​​feature vector.

[0047] Specifically, such as Figure 2 As shown, this embodiment first selects anchor frame c and the corresponding text description, for example, it can be randomly selected from A frame is selected as the anchor frame and input into the CLIP model (Contrastive Language-Image Pre-training, a multimodal pre-training model). For example, the description text vector obtained by the CLIP text encoder is , the dimension is , b is the batch size; the vector obtained by the image encoder is for ,after After a MLP (multi-layer perceptron) feature mapping, it is mapped to Same dimensions, then Perform vector concatenation in the last dimension to obtain a vector dimension of , and then through one-dimensional convolution and normalization, the features are mapped to , recorded as It should be noted that a weight needs to be initialized for each stage in the network. (The number of stages here is different from that of the Vincent video model and feature injector, such as Figure 2 As shown in Figure 8), .Then and The target vector can be obtained through the modal content interaction layer, that is, the modal interaction vector :

[0048] ;

[0049] Then, the vector generated by each stage of the Wensheng video model is combined with the input random noise (such as ModelScope), you can calculate the fusion features :

[0050] ;

[0051] Among them, as shown in the above formula, in this embodiment, The last four characteristics.

[0052] Step S14: Generate a new current first video based on the fusion features, and jump to the step of filtering out a preset number of target video frames from the current first video until the sum of the number of frames of each first video meets the preset frame number condition, and then generate a corresponding second video based on each first video.

[0053] In this embodiment, a new current first video is generated based on the fusion feature, and then Figure 2 As shown, jump to the step of filtering out a preset number of target video frames from the current first video until the sum of the number of frames of each first video meets the preset frame number condition, and then generate the corresponding second video based on each first video. After the Vincent video model performs a forward pass and completes a denoising, repeat the above steps of selecting the target video frame and fusion feature denoising for a total of T times (T can be an empirical value and is not specifically limited. For example, in this embodiment, T=5 is taken), and T denoising is completed. Finally, the next video content with a length of 32 frames is obtained through the decoder. After that, the obtained video content is used as the new initial video content again, and the steps of the last new video are repeated until the video content length that meets the user's needs is generated to obtain the corresponding long video. That is to say, in the process of generating a new current first video based on the fusion feature in this embodiment, if Figure 2 As shown, the random enhanced hybrid inference module can be used to improve the spatial resolution of the long video content generated in the previous step to complete the generation of high-quality video content. Specifically, the current fusion feature can be input into the preset image encoder, and the feature vector output by the preset image encoder based on the current fusion feature is used as the new first feature vector, and then the step of generating the corresponding conditional feature based on the first feature vector and the second feature vector is jumped to, until the current fusion feature meets the preset denoising condition, then the new current first video is generated based on the current fusion feature. It can be understood that the above-mentioned preset denoising condition can be the number of times the fusion feature is re-input into the preset image encoder.

[0054] It should be pointed out that the first video in this embodiment is a short video whose video length meets the preset short-length determination conditions, and the second video is a long video whose video length meets the preset long-length determination conditions. That is to say, in this embodiment, the first video and the second video are both videos that meet the corresponding duration or frame number conditions, and the above conditions can be adjusted according to the actual situation of the model and are not specifically limited to videos of a certain length. It can be understood that since the second video is generated by multiple first videos, the duration or number of frames of the second video is greater than that of the first video. For example, the first video in this embodiment can be a video with a frame number of 32 frames, while the second video is a video with a frame number greater than 1000 frames or a video with a duration greater than 1 minute.

[0055] Specifically, the second video can be segmented based on a preset video segmentation length to obtain a number of segmented videos; wherein there are a target number of repeated video frames between adjacent segmented videos; then Gaussian noise random sampling is performed on each segmented video to obtain a sampled video, and the sampled video is denoised based on the second eigenvector using a preset video diffusion super-resolution model, and an optimized second video is obtained based on the denoised sampled video. In the process of denoising the sampled video, the current sampled video can be denoised first to obtain the first latent vector , and denoise the next sampled video of the current sampled video to obtain the second latent vector , then determine the repeated video frames between the current sampled video and the next sampled video, and determine the third latent vector corresponding to the repeated video frames in the first latent vector and the second latent vector and the fourth latent vector , then based on the third latent vector and the fourth latent vector, the current sampled video is denoised.

[0056] Specifically, such as Figure 4 As shown, the second video can be evenly divided first, each segment contains L frames (32 frames in this embodiment), and adjacent segments have M frames repeated (in this embodiment, Figure 4 8 frames are shown), and then the corresponding Gaussian noise random sampling is performed on each video segment:

[0057] ;

[0058] ;

[0059] Afterwards Combined with Fi, a video diffusion super-resolution model (such as MS-Vid2Vid-XL) is used for denoising. The denoising step number is t=100 steps, and the diffusion step number of the whole process is T=1000 steps. Then the denoised latent vector is obtained. It is understandable that the reason for t=100 steps here is that the early steps of the denoising process are mainly to maintain the overall content structure of the video, while the subsequent denoising process gradually increases the details of the video content.

[0060] like Figure 4 As shown, the random mixing strategy is then executed to perform t-step denoising operations on the i-1th video clip and the i-th video clip respectively, and the denoised latent vector is obtained. and , and select the denoising latent vector of the corresponding part of its common coverage video frame and .in, Covering the Smooth temporal content information from the start frame of each video clip to the common overlay video frame; Covering the The video clips share the smooth temporal content information from the video frame to their end frame.

[0061] When denoising the current sampled video, a random video frame position can be determined from the repeated video frames, and based on the random video frame position, a corresponding first video frame in the current sampled video can be determined. Then, based on the video frame position next to the random video frame position, a corresponding second video frame in the next sampled video can be determined. A first to-be-spliced ​​vector corresponding to the first video frame is then determined from the third latent vector, and a second to-be-spliced ​​vector corresponding to the second video frame is determined from the fourth latent vector. The first to-be-spliced ​​vector and the second to-be-spliced ​​vector are then spliced ​​together to obtain a splicing vector. The current sampled video is denoised based on the splicing vector, and the next sampled video is used as the new current sampled video. The process then jumps to the step of denoising the current sampled video, and so on until denoising of each sampled video is completed.

[0062] That is to say, in order to solve the inconsistency of the generated content of the two video clips, this embodiment can use the denoised latent vector and For example, randomly pick a number m from [0, M], and then and Select the two adjacent frames of hidden vectors starting with the number (for example Select the mth latent vector, Select the m+1th latent vector) and perform vector concatenation, so that the denoised latent vectors between segments can interact with each other and reduce the content inconsistency generated by the super-resolution model in the video sub-segments. Then perform the The next step of video segment denoising is Denoising subsequent frames of a video segment , the corresponding denoising latent vector needs to be multiplied by the denoising probability, which is specifically , by performing 9000 denoising operations, the The super-resolution denoising process for each segment is performed. By performing the above denoising process on all video segments, the super-resolution denoising process for the entire long video can be completed. In this way, the random enhancement hybrid inference module can enhance the spatial resolution of the generated long video content, completing the generation of high-quality video content.

[0063] In this embodiment, a regression-based long video generation system is designed to address the problem that current video models cannot produce high-quality long videos. This system uses an autoregressive network to generate long video content. The autoregressive network module mainly includes a historical frame conditional attention network module, a long-term memory perception module, and an autoregressive video random enhancement hybrid module. The historical frame conditional attention network module captures the visual features of historical frames, maintaining the continuity of future generated video frames; the long-term memory perception module captures high-level scene semantics from earlier history, preventing forgetting during future video frame generation; and the autoregressive video random enhancement hybrid module processes video clips with visually repetitive content, maintaining the consistency and smoothness of the generated long video.

[0064] Based on the previous embodiment, it can be seen that the present application can use the feature injector to output the features of the target video frame, and determine the fusion features after generating the corresponding target vector based on the anchor frame and the description text, thereby generating future video frames. It should be pointed out that when a large model with longer visual tokens is currently used to generate a long video model, such a large model not only requires more data, but also requires a very large computing power cluster. Therefore, due to the limitation of computing power, it is difficult to generate a high-quality long video model. Therefore, in response to the problem that the current Wensheng video model cannot produce high-quality long videos with lower computing power requirements, a regression-based long video generation system is designed in combination with the previous embodiment. Next, the acquisition process of the initial target video frame will be described in detail in this embodiment. See Figure 5 As shown, the embodiment of the present application provides a specific video frame acquisition method, including:

[0065] Step S21: Generate an initial key frame corresponding to the description text using a preset key frame generation model.

[0066] In this embodiment, it is understood that generally, transitioning from a Vincent image model to a Vincent video model requires designing a temporal attention module based on the original model. However, this additional temporal attention module requires training on a large number of specific video-text pair datasets to obtain a good Vincent video model. Video-text pairs consume significantly more computing power than image-text pairs. The short video generation model proposed in this solution utilizes a novel regression-based generation method. On the one hand, this method only requires the use of a traditional Vincent graph network model, eliminating the need to train a temporal attention module, significantly reducing computing power consumption. On the other hand, this novel regression-based generation method utilizes a layered, parallel output method for video frames, significantly improving generation efficiency.

[0067] It should be noted that the short video generation network model in this embodiment mainly includes a keyframe text graph model (i.e., a keyframe generation model) and a keyframe interpolation model, both of which adopt a generation method based on grid images. The input of the keyframe text graph model is 512*512 Gaussian noise and text input, and the output is four keyframe images, each of which is 256*256, where the temporal order of the keyframes is from left to right and from top to bottom. Keyframe interpolation can be regarded as an image editing problem. The keyframe images are placed in the upper left and lower right corners, and Gaussian noise is placed in the upper right and lower left corners. The image size is also 256*256. Then, a video frame between the two keyframes is generated based on the text input. It should also be noted that for crawling video data from the Internet, it is first necessary to create a keyframe selection model. The criteria for keyframe selection are scene change and motion change. Correspondingly, in this embodiment, the key frames of the video for text description can be extracted based on the key frame selection network of the large model. Each video V will obtain the corresponding video description C and key frame sequence S. After obtaining the corresponding data, the model training of key frame generation and key frame insertion can be performed based on the existing text-based graph model.

[0068] Specifically, in this embodiment, a preset key frame generation model can be used to first generate initial key frames corresponding to the description text. Specifically, a preset key frame extraction model can be used to extract corresponding key frames from a preset video, and the key frames can be spliced ​​to obtain a corresponding mixed image. Then, the initial key frame generation model can be used to generate a Markov chain latent vector, and the Markov chain latent vector can be added to the mixed image. After that, the cross-attention layer in the initial key frame generation model is used to denoise the mixed image based on the Markov chain latent vector. After obtaining the denoised image, the initial key frame generation model is trained based on the denoised image and the mixed image until the current loss value of the initial key frame generation model meets the first preset training condition, and the current initial key frame generation model is used as the preset key frame generation model.

[0069] like Figure 6As shown in Figure 2, the key frame generation model is a two-dimensional convolution diffusion model. In the key frame generation diffusion model network, the model adopts an encoder-decoder architecture, which is mainly composed of an encoder-decoder, a cross-attention downsampling layer, and a cross-attention upsampling layer. Among them, the network structure of each cross-attention layer is as follows: Figure 7 As shown in the figure, downsampling is a two-fold downsampling of the feature map, and upsampling is a 2-fold upsampling of the feature. The network structure of the upsampling network and the downsampling network is the same, mainly consisting of a residual module, a self-attention mechanism module, a feedforward network module and a cross-attention module. The residual module adopts the basic network module of the residual network ResNet.

[0070] When training the initial key frame generation model, first give four key frames in the order from left to right and from top to bottom. A mixed image composed of a combination of , then the diffusion model first generates a series of Markov chain hidden vectors , this series of latent vectors is gradually transferred to the original image Add Gaussian noise. The formula for adding Gaussian noise is as follows:

[0071] ;

[0072] Based on latent vector The specific formula of the denoising process is as follows:

[0073] ;

[0074] Specifically, in the process of generating images, the initial From Gaussian noise Random sampling, gradually reducing noise to obtain a real image It is understandable that in actual practice, the deviation of the predicted noise rather than the real image is usually used as the specific optimization training function. The specific function is as follows:

[0075] ;

[0076] in is from Gaussian noise Random sampling, It is the deviation obtained by neural network inference of the noisy image at time step t and time t under input condition c. are parameters learned by the neural network, and c represents various conditional inputs, which in this embodiment refers to text descriptions.

[0077] The loss function is then used to judge the training. The specific loss function is a loss function that combines the deviation prediction and the actual value prediction as the regularization term. The overall loss function is as follows:

[0078] ;

[0079] At this time Indicates the number of parameters that can be learned at present, The latter term acts as a regular term to reduce the risk of overfitting.

[0080] Step S22: Using a preset key frame interpolation model to interpolate based on the initial key frame to obtain the current first video; the preset key frame generation model and the preset key frame interpolation model are models constructed based on a two-dimensional convolution diffusion model.

[0081] In this embodiment, a preset key frame interpolation model can be used to interpolate based on the initial key frame to obtain the current first video. Wherein, as in the previous step, the preset key frame generation model and the preset key frame interpolation model are both models constructed based on the two-dimensional convolution diffusion model. The training process of the key frame interpolation model specifically includes: first, constructing an initial key frame interpolation model based on the network architecture and model weights of the preset key frame generation model, and generating corresponding Gaussian noise according to the resolution of the key frame, constructing an initial input image based on the key frame and Gaussian noise, and then inputting the initial input image into the initial key frame interpolation model to interpolate between key frames to obtain corresponding inserted video frames, and then training the initial key frame interpolation model based on the inserted video frames and key frames until the current loss value of the initial key frame interpolation model meets the second preset training condition, and the current initial key frame interpolation model is used as the preset key frame interpolation model.

[0082] Specifically, such as Figure 8 As shown in Figure 1, the keyframe insertion model changes the input of the keyframe generation model, and the overall network architecture remains the same. Therefore, the keyframe generation model can be loaded with the initial weights of the network model, and the training method and loss function are consistent with the keyframe generation model. Correspondingly, the process of inferring the initial video frame from the keyframe text graph model and the insertion model is as follows: Figure 9 As shown in the figure: For user input text, firstly, the key frames based on scene conversion and motion conversion related to the text input are generated according to the key frame text graph model, for example Figure 9In the example, the images of the 1st, 10th, 21st and 30th frames are generated, and then the key frame interpolation model is used to interpolate adjacent key frames in a recursive and parallel manner. For example, adjacent key frames are placed in the upper left corner and lower right corner with a resolution of 256*256 respectively, and then Gaussian noise is added to the lower left corner and upper right corner respectively, thus forming an initial input of 512*512. The initial input mixed with key frames and Gaussian noise is then input into the key frame interpolation model for interpolation to generate intermediate video frames. For example, the 1st and 10th frames generate intermediate frames of the 4th and 7th frames, and the 21st and 30th frames generate intermediate frames of the 24th and 27th frames. And it can be understood that if you continue to generate high frame rate videos, you can continue to call the key frame interpolation model recursively and in parallel.

[0083] Before using the preset key frame extraction model to extract the corresponding key frames from the preset video, a training video can also be obtained, the corresponding video content text can be generated according to the training video, and the corresponding positive samples and negative samples can be constructed based on the training video and the video content text to construct a training set based on the positive samples and negative samples, and then the image encoder in the initial key frame extraction model can be used to convert the training video in the training set into the corresponding video frame vector, and the key frame probability corresponding to each video frame of the training video can be determined according to the video frame vector, so as to train the initial key frame extraction model based on the key frame probability and the training video, until the current loss value of the initial key frame extraction model meets the third preset training condition, and the current initial key frame extraction model is used as the preset key frame extraction model.

[0084] That is to say, if Figure 10 As shown in the figure, when training the video key frame extraction model, first obtain the Internet video set V (V video limit length is within 10s), and then each video , input into the Video Caption model (such as GPT4V API) to obtain a detailed description of the video content , and get the data pair , then the data Construct positive and negative samples, the positive sample is , negative samples have the following methods: The first is to Randomly replace the objects or attributes in ; The second is to retrieve from the collection Similar descriptions ,constitute Then control the ratio of positive and negative samples to 1:2, obtain the training set D, and follow Figure 10The network model shown uses D to train the key frame extraction network. Specifically, in this embodiment, the image encoder uses ViT, and the key frame network (the network structure is Bert-Base, and the initial weights loaded are also Bert-Base) is the network weight that needs to be trained. The video frame passes through the image encoder to obtain the frame vector . Among them, the query vector is the randomized initial parameter q, q and the frame vector are input into the key frame network to obtain the solidified vector h. It should be pointed out that the role of q is to fix it to the same time dimension regardless of the number n. For example, the time dimension in this embodiment is 100. Thus, h is combined with the question prompt and the video description and input into the large language model (such as vicuna-13b). The large language model extracts the key frames that match the video description, and then inputs the key frame features into the answer network, wherein the answer network and the key frame network share weights. In this way, the probability that the large language model outputs yes can be used as the probability of whether the video frame is a key frame, and the training loss is the cross entropy loss used by the large language model to predict the next token task. After completing the key frame network training, Inference can get the key frames of each video .

[0085] Step S23: taking the last frame of the current first video as the starting frame, and sequentially screening a preset number of video frames in the current first video based on the starting frame to obtain a target video frame.

[0086] In this embodiment, the last frame of the current first video is used as the starting frame, and a preset number of video frames in the current first video are sequentially filtered based on the starting frame to obtain the target video frame. For example, in this embodiment, the last 16 frames of the 32-frame current first video can be extracted as the target video frame.

[0087] Through the above technical solution, in this embodiment, a fast video content generation model based on the literary graph model is designed to generate short video sequences to address the problem that the current literary video model cannot produce high-quality long videos with low computing power requirements. The model adopts the traditional two-dimensional convolutional diffusion model UNet instead of the three-dimensional convolutional diffusion model with higher computing power consumption. At the same time, the model does not need to train the temporal network module, and adopts a regression-type space-for-time reasoning method, which can perform parallel calculations and realize efficient and fast reasoning of short video content.

[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0089] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned video generation method embodiments.

[0090] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned video generation method embodiments when running.

[0091] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0092] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned video generation method embodiments are implemented.

[0093] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned video generation method embodiments are implemented.

[0094] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0095] The above is a detailed introduction to the video generation method, device, storage medium, and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core concept of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, various improvements and modifications may be made to the present application, and such improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A video generation method, characterized in that: include: Filtering a preset number of target video frames from the current first video, and outputting a plurality of first feature vectors corresponding to the target video frames using a preset feature injector; Generate a plurality of second feature vectors based on preset noise and a description text corresponding to the current first video using a preset Wensheng video model, and generate corresponding conditional features based on the first feature vector and the second feature vector; Determining an anchor frame in the current first video, generating a corresponding target vector according to the anchor frame and the description text, and determining a corresponding fusion feature based on the target vector, the first feature vector, and the conditional feature; A new current first video is generated based on the fusion features, and the process jumps to the step of filtering out a preset number of target video frames from the current first video until the sum of the number of frames of each first video meets the preset frame number condition, and then a corresponding second video is generated based on each first video.

2. The video generation method according to claim 1, characterized in that The step of selecting a preset number of target video frames from the current first video includes: generating an initial key frame corresponding to the description text using a preset key frame generation model, and performing frame insertion based on the initial key frame using a preset key frame insertion model to obtain the current first video; the preset key frame generation model and the preset key frame insertion model are models constructed based on a two-dimensional convolution diffusion model; The last frame of the current first video is used as a starting frame, and based on the starting frame, the preset number of video frames in the current first video are sequentially screened to obtain the target video frame.

3. The video generation method according to claim 2, characterized in that Before filtering out a preset number of target video frames from the current first video, the method further includes: Extracting corresponding key frames from a preset video using a preset key frame extraction model, and splicing the key frames to obtain a corresponding mixed image; generating a Markov chain latent vector using an initial key frame generation model, and adding the Markov chain latent vector to the mixed image; Denoising the mixed image based on the Markov chain latent vector using a cross attention layer in the initial key frame generation model to obtain a denoised image; The initial key frame generation model is trained based on the denoised image and the mixed image until a current loss value of the initial key frame generation model meets a first preset training condition, and the current initial key frame generation model is used as the preset key frame generation model.

4. The video generation method according to claim 3, wherein: Before filtering out a preset number of target video frames from the current first video, the method further includes: Constructing an initial key frame insertion model based on the network architecture and model weights of the preset key frame generation model; Generate corresponding Gaussian noise according to the resolution of the key frame, and construct an initial input image based on the key frame and the Gaussian noise; Inputting the initial input image into the initial key frame interpolation model so as to interpolate frames between the key frames to obtain corresponding interpolated video frames; The initial key frame interpolation model is trained based on the inserted video frame and the key frame until a current loss value of the initial key frame interpolation model meets a second preset training condition, and the current initial key frame interpolation model is used as the preset key frame interpolation model.

5. The video generation method according to claim 4, characterized in that: Before extracting corresponding key frames from a preset video using a preset key frame extraction model, the method further includes: Obtaining a training video, and generating corresponding video content text based on the training video; Constructing corresponding positive samples and negative samples based on the training video and the video content text, so as to construct a training set according to the positive samples and the negative samples; Converting the training videos in the training set into corresponding video frame vectors using an image encoder in an initial key frame extraction model; Determining a key frame probability corresponding to each video frame of the training video according to the video frame vector; The initial key frame extraction model is trained based on the key frame probability and the training video until a current loss value of the initial key frame extraction model meets a third preset training condition, and the current initial key frame extraction model is used as the preset key frame extraction model.

6. The video generation method according to claim 1, wherein: The step of outputting a plurality of first feature vectors corresponding to the target video frame by using a preset feature injector includes: Processing the target video frame using a preset image encoder and a convolutional layer, and adding Gaussian noise to the processed target video frame to obtain a noisy target video frame; The noisy target video frame is input into the preset feature injector to output the first feature vector corresponding to the target video frame of several levels; the levels are divided based on the network structure of the preset image encoder, and the network structure and initialization parameters of the preset image encoder are the same as those of the preset feature injector.

7. The video generation method according to claim 6, characterized in that: Generating corresponding conditional features based on the first feature vector and the second feature vector includes: Normalizing the second feature vector using a preset spatiotemporal normalization layer, and processing the normalized second feature vector using a first preset fully connected layer; The first feature vector and the processed second feature vector are fused through a preset temporal attention layer, and the fused feature vector is processed using a second preset fully connected layer to obtain the conditional feature.

8. The video generation method according to claim 7, characterized in that: The process of generating a new current first video based on the fusion feature further includes: Inputting the current fusion feature into the preset image encoder; The feature vector output by the preset image encoder based on the current fusion feature is used as the new first feature vector, and the process jumps to the step of generating corresponding conditional features based on the first feature vector and the second feature vector until the current fusion feature meets the preset denoising condition, and then the new current first video is generated based on the current fusion feature.

9. The video generation method according to claim 1, wherein: Generating a corresponding target vector according to the anchor frame and the description text includes: Encoding the description text using a text encoder of a preset multimodal training model to obtain a corresponding text feature vector; Encoding the anchor frame using the image encoder of the preset multimodal training model to obtain a corresponding image feature vector; Using a preset multi-layer perceptron, feature mapping is performed on the image feature vector according to the vector dimension of the text feature vector, and the text feature vector and the mapped image feature vector are spliced ​​to obtain a spliced ​​feature vector; The weights corresponding to the various levels of the preset Vincent video model are determined, and the corresponding target vector is generated according to the weights and the spliced ​​feature vector.

10. The video generation method according to any one of claims 1 to 9, characterized in that: After generating the corresponding second video based on each first video, the method further includes: Segmenting the second video based on a preset video segmentation length to obtain a plurality of segmented videos; wherein a target number of repeated video frames exist between adjacent segmented videos; Performing Gaussian noise random sampling on each of the segmented videos to obtain a sampled video; The sampled video is denoised according to the second eigenvector using a preset video diffusion super-resolution model, and an optimized second video is obtained based on the denoised sampled video.

11. The video generation method according to claim 10, characterized in that: The denoising of the sampled video includes: Denoising the current sampled video to obtain a first latent vector, and denoising the next sampled video of the current sampled video to obtain a second latent vector; Determining the repeated video frame between the current sampled video and the next sampled video, and determining a third latent vector corresponding to the repeated video frame from the first latent vector, and determining a fourth latent vector corresponding to the repeated video frame from the second latent vector; Based on the third latent vector and the fourth latent vector, denoising the current sampled video again; The next sampled video is used as a new current sampled video, and the process jumps to the step of denoising the current sampled video, so as to denoise each of the sampled videos.

12. The video generation method according to claim 11, characterized in that: The denoising the current sampled video again based on the third latent vector and the fourth latent vector includes: Determining a random video frame position from the repeated video frames, and determining a corresponding first video frame in the current sampled video based on the random video frame position; Determining a corresponding second video frame in the next sampled video according to a video frame position next to the random video frame position; Determine a first to-be-spliced ​​vector corresponding to the first video frame from the third latent vector, and determine a second to-be-spliced ​​vector corresponding to the second video frame from the fourth latent vector; The first vector to be spliced ​​and the second vector to be spliced ​​are spliced ​​to obtain a spliced ​​vector, and denoising is performed on the current sampled video according to the spliced ​​vector.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the video generation method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the video generation method according to any one of claims 1 to 12 are implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the video generation method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Long video understanding method based on iterative hierarchical key frame selection

    CN119785258A

  • Video generation method and device, computer program product and electronic equipment

    CN120017929A