Video generation method and device, equipment and storage medium

By integrating multiple control conditions in the video generation model, the problems of low quality and low efficiency of image and video generation in the prior art are solved, and high-quality and efficient video generation are achieved.

CN120201259APending Publication Date: 2025-06-24BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510266150.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When generating controllable images and videos, discrete image representations lead to information loss, resulting in low quality of generated images or videos and low generation efficiency.

Method used

By using multiple signal processing modules in the video generation model to process multiple data signals, multiple long sequence listings are generated, and multiple control conditions are integrated in the sequence dimension, so that all control conditions can interact in the same learning process.

Benefits of technology

The quality of the generated video is improved, the generation efficiency is improved, and the information integrity and accuracy in the video generation process is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201259A_ABST
    Figure CN120201259A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: processing various data signals through a plurality of signal processing modules in a video generation model to obtain a plurality of long sequence representations; processing the plurality of long sequence representations through a video diffusion model in a video generation model to obtain target video features; and decoding the target video feature to obtain a target video. According to the scheme, through comprehensive processing of various data signals, video details and features can be accurately captured, so that high-quality target video features are generated, the target video is further obtained, the quality of the generated video is improved, and the video generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a video generation method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technologies, an important research direction in the fields of computer vision and artificial intelligence is multi-modal controllable image and video generation. And the future development direction of controllable image and video generation is higher quality, stronger generalization ability, and higher computational efficiency.

[0003] Currently, a common controllable image and video generation solution is to use discrete image representations, discretize images into tokens, and regard them as part of language vocabulary. However, discrete image tokens need to be processed through two stages of encoding and decoding. After encoding and then decoding the images, huge losses will occur, resulting in information loss and low quality and low generation efficiency of the generated images or videos. Summary of the Invention

[0004] The present disclosure provides a video generation method, apparatus, device, and storage medium. Compared with the prior art, this solution integrates multiple control conditions in the sequence dimension, enables all control conditions to interact in the same learning process, improves the quality of the generated videos, and thus improves the generation efficiency.

[0005] According to one aspect of the embodiments of the present disclosure, a video generation method is provided, characterized in that the method includes:

[0006] Processing multiple data signals respectively through multiple signal processing modules in a video generation model to obtain multiple long sequence representations. The multiple signal processing modules are used to process the input video or image into corresponding long sequence representations, and each long sequence representation corresponds to one data signal. The multiple data signals are used to provide at least one of noise information, camera trajectory information, depth information, and object information;

[0007] Processing the multiple long sequence representations through a video diffusion model in the video generation model to obtain target video features. The video diffusion model is used to output video features based on the input data;

[0008] Decoding the target video features to obtain a target video.

[0009] According to another aspect of the embodiments of the present disclosure, a video generation apparatus is provided, and the apparatus includes:

[0010] A first processing unit, configured to process multiple types of data signals respectively through multiple signal processing modules in a video generation model to obtain multiple long sequence representations. The multiple signal processing modules are used to process an input video or image into corresponding long sequence representations, and each long sequence representation corresponds to one type of data signal. The multiple types of data signals are used to provide at least one of noise information, camera trajectory information, depth information, and object information;

[0011] A second processing unit, configured to process the multiple long sequence representations through a video diffusion model in the video generation model to obtain target video features. The video diffusion model is used to output video features based on input data;

[0012] A decoding unit, configured to decode the target video features to obtain a target video.

[0013] In some embodiments, the first processing unit is configured to, for a depth video among the multiple types of data signals, input the depth video into a first signal processing module in the video generation model. The depth video includes the depth information; compress the depth video in time series and space through a variational autoencoder in the first signal processing module to obtain a depth spatio-temporal vector; slice the depth spatio-temporal vector in the spatial dimension through the first signal processing module to obtain multiple depth vector blocks; splice the multiple depth vector blocks in the sequence dimension through the first signal processing module to obtain a depth vector sequence; map the depth vector sequence block to a latent space through a convolutional layer in the first signal processing module to obtain a first long sequence representation, and the first long sequence representation is the long sequence representation of the depth video.

[0014] In some embodiments, the first processing unit is configured to, for a reference image among the multiple types of data signals, input the reference image into a second signal processing module in the video generation model. The reference image includes the object information; compress the reference image in space through a variational autoencoder in the second signal processing module to obtain an image space vector; slice the image space vector in the spatial dimension through the second signal processing module to obtain multiple image vector blocks; splice the multiple image vector blocks in the sequence dimension through the second signal processing module to obtain an image vector sequence; map the image vector sequence block to a latent space through a convolutional layer in the second signal processing module to obtain a second long sequence representation, and the second long sequence representation is the long sequence representation of the reference image.

[0015] In some embodiments, the first processing unit is configured to, for the camera trajectory video among the multiple data signals, input the camera trajectory video into a third signal processing module in the video generation model, where the camera trajectory video includes the camera trajectory information; through the third signal processing module, perform Plücker embedding on the camera trajectory video to obtain a camera trajectory vector; through the third signal processing module, perform slicing on the camera trajectory vector in the spatial dimension to obtain a plurality of camera trajectory vector blocks; through the third signal processing module, splice the plurality of camera trajectory vector blocks in the sequence dimension to obtain a camera trajectory vector sequence; through a convolutional layer in the third signal processing module, map the camera trajectory vector sequence block to a latent space to obtain a third long sequence representation, and the third long sequence representation is the long sequence representation of the camera trajectory video.

[0016] In some embodiments, the second processing unit is configured to splice the plurality of long sequence representations in the sequence direction to obtain a multi-modal fusion feature; input the multi-modal fusion feature into the video diffusion model; inject text features into the video diffusion model through a cross-attention mechanism, where the text features are used to provide text prompt information; through the video diffusion model in the video generation model, process the multi-modal fusion feature and the text features based on a full attention mechanism to obtain the target video feature.

[0017] In some embodiments, the device further includes:

[0018] A third processing unit, configured to perform feature extraction on a text prompt signal through a text encoder in the video generation model to obtain the text features.

[0019] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes:

[0020] One or more processors;

[0021] A memory for storing executable program code of the processor;

[0022] Wherein, the processor is configured to execute the program code to implement the above video generation method.

[0023] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above video generation method.

[0024] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the above video generation method.

[0025] Embodiments of the present disclosure provide a video generation solution. Multiple signal processing modules in a video generation model process various data signals to generate multiple long-sequence representations. The above-mentioned multiple long-sequence representations cover at least one type of information such as noise, camera trajectory, depth, and objects. By integrating various control conditions in the sequence dimension, it not only provides a rich and key data source for the video diffusion model, enabling the video diffusion model to understand videos or images from multiple dimensions, but also helps to accurately capture video details and features through the comprehensive processing of various data signals, thereby generating high-quality target video features, and then obtaining the target video, improving the quality of the generated video and the video generation efficiency.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0028] Figure 1 is a schematic diagram of an implementation environment of a video generation method shown according to an exemplary embodiment.

[0029] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment.

[0030] Figure 3 is a flowchart of another video generation method shown according to an exemplary embodiment.

[0031] Figure 4 is a schematic diagram of the structure of a video diffusion model provided according to an exemplary embodiment.

[0032] Figure 5 is a block diagram of a video generation device shown according to an exemplary embodiment.

[0033] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0035] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0036] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, various data signals involved in the present disclosure are obtained under full authorization.

[0037] Figure 1 It is a schematic diagram of an implementation environment of a video generation method shown according to an exemplary embodiment. Refer to Figure 1 , and the implementation environment specifically includes: an electronic device 101 and a server 102. The electronic device 101 can be connected to the server 102 through a wireless network or a wired network.

[0038] The electronic device 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop portable computer. An application program can be installed and run on the electronic device 101, and the application program is used to generate a video according to various data signals input by the user. The application program is associated with the server 102, and the server 102 provides background services to the electronic device 101.

[0039] The electronic device 101 can generally refer to one of multiple electronic devices, and the electronic device 101 is used as an example in this embodiment. Those skilled in the art can know that the number of the above-mentioned electronic devices can be more or less. For example, the above-mentioned electronic devices can be several, or the above-mentioned electronic devices can be dozens or hundreds, or more, and the present disclosure embodiment does not limit the number and type of the electronic devices.

[0040] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above servers may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may further include other functional servers to provide more comprehensive and diverse services. In some embodiments, the server 102 undertakes the main computing work, and the electronic device 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the electronic device 101 undertakes the main computing work; or, the server 102 and the electronic device 101 adopt a distributed computing architecture for collaborative computing. The server 102 may be connected to the electronic device 101 and other electronic devices through a wireless network or a wired network. Optionally, the number of the above servers may be more or less, and the embodiments of the present disclosure do not limit this.

[0041] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 2 shown, the method is executed by an electronic device and includes the following steps:

[0042] In step S201, multiple data signals are respectively processed by multiple signal processing modules in the video generation model to obtain multiple long sequence characterizations.

[0043] In the embodiments of the present disclosure, the video generation model is used to generate realistic, coherent, and video with specific content. Correspondingly, the video generation model requires a large amount of rich and accurate data signals as input, so as to be able to learn and understand sufficient information to meet the generation requirements of the model for complex video content, thereby generating high-quality videos. Optionally, the multiple data signals include at least one of a noise video, a camera trajectory video, a depth video, and a reference image. Correspondingly, the multiple data signals are used to provide at least one of noise information, camera trajectory information, depth information, and object information.

[0044] Among them, the noise video is used to provide noise information. This noise video is not a completely meaningless interference signal, but a random or pseudo-random signal with a specific role in video generation. The noise video can simulate various uncertainties and minute changes existing in the real world, adding a sense of reality and diversity to the generated video.

[0045] The camera trajectory video is used to provide camera trajectory information. This camera trajectory information describes the changes in the position, orientation, and movement path of the camera over time during video shooting. The camera trajectory information usually includes movement parameters such as the translation and rotation of the camera, as well as the spatial positions at different time points.

[0046] Depth videos are used to provide depth information. The depth information is used to represent the distance relationship between each object in the scene and the camera. The depth information can be the depth value corresponding to each pixel point, or a description of the depth levels of different objects in the scene. The depth information can assist the model in understanding the three-dimensional structure of the scene.

[0047] Reference images are used to provide object information. The object information is used to represent attributes such as the category, shape, size, color, and material of various objects in the video scene, as well as the mutual relationships between objects, such as the position relationship and the interaction relationship. Optionally, the reference image is two people, or one person and a house, and the embodiments of the present disclosure do not limit this.

[0048] The video generation model includes multiple signal processing modules. The multiple signal processing modules are used to process the input video or image into corresponding long sequence representations. Among them, different signal processing modules are used to process different types of data signals.

[0049] Among them, each long sequence representation corresponds to a type of data signal. Each long sequence representation is a highly integrated and abstract representation of a type of data signal. The long sequence representation extracts, compresses, and encodes various detailed information in the original data signal, and represents the key features and change rules of the data signal in a more compact and meaningful way. The multiple long sequence representations provide a basis for multi-modal fusion in the video generation model.

[0050] In step S202, the multiple long sequence representations are processed through the video diffusion model in the video generation model to obtain target video features.

[0051] In the embodiments of the present disclosure, the above-mentioned multiple long sequence representations contain key information extracted from different data signals, such as the noise information, camera trajectory information, depth information, and object information mentioned above. These long sequence representations provide multi-dimensional basic information about the video content for the video diffusion model, helping the model understand various features and dynamic changes of the video. The video diffusion model is used to output video features based on the input data.

[0052] In the reverse denoising process of the video diffusion model, the multiple long sequence representations are used as conditional information to guide the denoising process. The video diffusion model can predict how to remove noise at each denoising step according to this conditional information, so as to gradually restore the target video features.

[0053] In step S203, the target video features are decoded to obtain the target video.

[0054] In the embodiments of the present disclosure, the decoder generally adopts a neural network architecture. Commonly, there are the transposed convolutional neural network (also known as the transposed convolution network) and the generator part in the generative adversarial network (GAN), etc.

[0055] Optionally, starting from the target video features, through multiple deconvolution or upsampling operations, the width and height of the feature map are gradually expanded until the resolution of the target video is reached. During the process of generating the video, the content of the video frames can be continuously adjusted to make them temporally coherent and avoid jumps or unnatural transitions. The decoded video frames may also need to undergo some post-processing operations, such as color correction, denoising, sharpening, etc., to improve the quality and visual effect of the video.

[0056] The embodiments of the present disclosure provide a video generation scheme. Multiple signal processing modules in the video generation model process multiple data signals to generate multiple long sequence representations. The above-mentioned multiple long sequence representations cover at least one type of information such as noise, camera trajectory, depth, and objects. By integrating multiple control conditions in the sequence dimension, it not only provides a rich and key data source for the video diffusion model, enabling the video diffusion model to understand videos or images from multiple dimensions, but also helps to accurately capture video details and features through the comprehensive processing of multiple data signals, thereby generating high-quality target video features and further obtaining the target video, improving the quality of the generated video and the video generation efficiency.

[0057] In some embodiments, multiple signal processing modules in the video generation model process multiple data signals respectively to obtain multiple long sequence representations, including:

[0058] For the depth video among the multiple data signals, the depth video is input into the first signal processing module in the video generation model, and the depth video includes depth information;

[0059] Through the variational autoencoder in the first signal processing module, the depth video is compressed in the temporal and spatial dimensions to obtain a depth spatio-temporal vector;

[0060] Through the first signal processing module, the depth spatio-temporal vector is sliced in the spatial dimension to obtain multiple depth vector blocks;

[0061] Through the first signal processing module, multiple depth vector blocks are concatenated in the sequence dimension to obtain a depth vector sequence;

[0062] Through the convolution layer in the first signal processing module, the depth vector sequence block is mapped to the latent space to obtain a first long sequence representation, which is a long sequence representation of the depth video.

[0063] In the disclosed embodiment,

[0064] In some embodiments, multiple signal processing modules in the video generation model process multiple data signals respectively to obtain multiple long sequence representations, including:

[0065] For a reference image among the plurality of data signals, inputting the reference image into a second signal processing module in the video generation model, the reference image including object information;

[0066] The reference image is spatially compressed by a variational autoencoder in the second signal processing module to obtain an image space vector;

[0067] The image space vector is segmented in the space dimension by a second signal processing module to obtain a plurality of image vector blocks;

[0068] Through the second signal processing module, multiple image vector blocks are spliced ​​in the sequence dimension to obtain an image vector sequence;

[0069] The image vector sequence blocks are mapped to the latent space through the convolution layer in the second signal processing module to obtain a second long sequence representation, which is a long sequence representation of the reference image.

[0070] In the disclosed embodiment, the reference image is spatially compressed by a variational autoencoder, which reduces the amount of image data and makes the model operation more efficient, so that a large number of reference images can be processed quickly. The image space vector is disassembled according to the spatial dimension by the segmentation operation, so that the model can focus on the object information in each area and mine the local details; the image information is integrated according to the sequence dimension by the splicing operation to ensure the coherence of the image information. The image vector sequence is further refined by the convolution layer, and the key features of the reference image are condensed. The above method improves the accuracy and richness of the video content, makes the generation results more in line with expectations, and improves the efficiency of video generation.

[0071] In some embodiments, multiple signal processing modules in the video generation model process multiple data signals respectively to obtain multiple long sequence representations, including:

[0072] For a camera trajectory video among the multiple data signals, input the camera trajectory video into a third signal processing module in the video generation model, where the camera trajectory video includes camera trajectory information;

[0073] Through the third signal processing module, the camera trajectory video is subjected to Plücker embedding to obtain a camera trajectory vector;

[0074] Through the third signal processing module, the camera trajectory vector is segmented in the spatial dimension to obtain multiple camera trajectory vector blocks;

[0075] Through the third signal processing module, multiple camera trajectory vector blocks are concatenated in the sequence dimension to obtain a camera trajectory vector sequence;

[0076] Through the convolutional layer in the third signal processing module, the camera trajectory vector sequence block is mapped to the latent space to obtain a third long sequence representation, and the third long sequence representation is the long sequence representation of the camera trajectory video.

[0077] In the embodiments of the present disclosure, the complex camera trajectory information is converted into vectors through Plücker embedding, effectively simplifying the data form, facilitating subsequent processing, and improving the calculation efficiency. By segmenting the camera trajectory vector in the spatial dimension, the camera motion details at different spatial positions can be focused on to capture local features. By concatenating the vectors in the sequence dimension, the coherence of the camera motion in time is ensured, and the entire trajectory is presented completely. The convolutional layer further mines the feature patterns in the camera trajectory vector sequence, maps it to the latent space, generates a third long sequence representation, and accurately extracts the key information of the camera trajectory. The third long sequence representation can provide an accurate camera motion reference for video generation, making the perspective switching and scene movement of the generated video more natural and smooth, highly fitting the actual shooting effect, greatly enhancing the realism and visual experience of the video, and improving the video generation efficiency.

[0078] In some embodiments, through the video diffusion model in the video generation model, multiple long sequence representations are processed to obtain target video features, including:

[0079] Multiple long sequence representations are concatenated in the sequence direction to obtain a multimodal fusion feature;

[0080] The multimodal fusion feature is input into the video diffusion model;

[0081] The text feature is injected into the video diffusion model through the cross-attention mechanism, and the text feature is used to provide text prompt information;

[0082] Through the video diffusion model in the video generation model, the multimodal fusion feature and the text feature are processed based on the full attention mechanism to obtain the target video feature.

[0083] In the embodiments of the present disclosure, by splicing multiple long sequence representations in the sequence direction, integrating multi-source information such as depth, images, and trajectories, a comprehensive basis is provided for video generation, enriching the video content dimension. By inputting the multi-modal fusion features into the video diffusion model, the model can perform diffusion and denoising based on multivariate information, generating video features that better fit complex scenarios. By injecting text features into the video diffusion model through the cross-attention mechanism, the generation direction is accurately guided to ensure that the video meets the user's semantic requirements. Through the video diffusion model, based on the full attention mechanism, the multi-modal fusion features and text features are processed, comprehensively considering feature interactions, not only capturing the associations within each modality but also exploring the connections between different modalities, enabling the elements of the video to cooperate and generating target video features with logical coherence and rich details, significantly improving the quality and accuracy of the generated video and enhancing the video generation efficiency.

[0084] In some embodiments, the method further includes:

[0085] Through the text encoder in the video generation model, feature extraction is performed on the text prompt signal to obtain text features.

[0086] In the embodiments of the present disclosure, by performing feature extraction on the text prompt signal through the text encoder in the video generation model to obtain text features, the natural language input by the user can be converted into key semantic features understandable by the model, accurately guiding the video generation direction and ensuring that the generated content meets the user's expectations.

[0087] The above Figure 2 shows a flowchart of a video generation method of the present disclosure. The video generation solution provided by the present disclosure will be further elaborated below. Figure 3 is a flowchart of another video generation method shown according to an exemplary embodiment. Refer to Figure 3 , which is executed by an electronic device and includes the following steps:

[0088] In step S301, multiple data signals are input into the video generation model.

[0089] In the embodiments of the present disclosure, the multiple data signals include at least one of a noise video, a camera trajectory video, a depth video, and a reference image. Among them, the noise video includes noise information. The camera trajectory video includes camera trajectory information. The depth video includes depth information. The reference image includes object information. Among them, the object indicated by the object information can be a person or an object, and the embodiments of the present disclosure do not limit this. Correspondingly, the multiple data signals are used to provide at least one of noise information, camera trajectory information, depth information, and object information.

[0090] The explanations of the noise video, the camera trajectory video, the depth video, and the reference image refer to the above step S201 and will not be elaborated here.

[0091] It should be noted that the video generation model includes multiple signal processing modules, a video diffusion model, a text encoder, and other components. The multiple signal processing modules are used to process the input video or image into corresponding long-sequence features. Among them, one signal processing module is used to process one type of data signal. Correspondingly, one long-sequence representation corresponds to one type of data signal.

[0092] The first signal processing module, the second signal processing module, the third signal processing module, the video diffusion model, and the text encoder are described below by way of example. Among them, the first signal processing module is used to process the input depth video. The second signal processing module is used to process the input reference image. The third signal processing module is used to process the input camera trajectory video. Since the video diffusion model also needs to input noise as the starting point for denoising, a signal processing module is used to process the noise video as the input of the video diffusion model. Optionally, there may be other data signals as control conditions in the process of generating the video, such as a segmentation map video and a human pose control video, etc., which will not be exemplified one by one in the embodiments of the present disclosure.

[0093] In step S302, the depth video in the multiple data signals is processed by the first signal processing module in the video generation model to obtain a first long-sequence representation, and the first long-sequence representation is the long-sequence representation of the depth video.

[0094] In the embodiments of the present disclosure, the depth video is a special type of video data. The depth video not only contains the color information (RGB) of a normal video, but also includes the distance information from each pixel point to the camera, that is, depth information. Among them, the depth information can endow the video scene with a three-dimensional structure, enabling the model to better understand and learn the spatial position relationship, front and back occlusion relationship, etc. between objects.

[0095] In some embodiments, a variational auto-encoder (VAE) is included in the first signal processing module. The depth video is compressed by the variational auto-encoder, then segmented and spliced, and finally mapped to the latent space through a convolutional layer to obtain the first long sequence representation. Correspondingly, first, for the depth video among multiple data signals, the depth video including depth information is input into the first signal processing module in the video generation model. Secondly, the depth video is compressed temporally and spatially by the variational auto-encoder in the first signal processing module to obtain a depth spatio-temporal vector. Thirdly, the depth spatio-temporal vector is segmented in the spatial dimension through the first signal processing module to obtain multiple depth vector blocks. Fourthly, multiple depth vector blocks are spliced in the sequence dimension through the first signal processing module to obtain a depth vector sequence. Finally, the depth vector sequence is mapped to the latent space through the convolutional layer in the first signal processing module to obtain the first long sequence representation, and the first long sequence representation is the long sequence representation of the depth video.

[0096] The following introduces the process of compressing the depth video temporally and spatially.

[0097] The variational auto-encoder is a generative model composed of an encoder and a decoder. The encoder maps the input data to the probability distribution of the latent space (latent space), usually assumed to be a Gaussian distribution; the decoder reconstructs the original input data from the samples in the latent space. In the processing of depth videos, the main role of the VAE is to compress the data, remove redundant information, and retain key features at the same time. The depth video is sequential data with a time dimension and contains multiple consecutive frames.

[0098] Temporally, the VAE can analyze the correlation and variation law between adjacent frames, and compress this temporal information into a lower-dimensional representation through learning. For example, taking a moving object in a depth video as an example, the VAE can capture the key features of the object's motion trajectory, and compared with the scheme of storing the detailed position information of the object in each frame, it can reduce the data volume.

[0099] In the spatial dimension, the VAE can process the depth information of each frame. Each pixel in the depth video has a corresponding depth value. The VAE can identify the pixels and regions that have an important impact on the overall depth structure, ignore some unimportant details, and compress the spatial information into a more compact representation. For example, in a complex three-dimensional scene, the VAE can extract the main contour of the scene and the general shape of the objects.

[0100] In the above way, the VAE can compress the depth video into a depth spatio-temporal vector, and the depth spatio-temporal vector can represent the features of the depth video in time and space in a more concise way.

[0101] The process of slicing the deep spatio-temporal vector in the spatial dimension is introduced below.

[0102] Slicing the deep spatio-temporal vector in the spatial dimension can divide the entire depth information into multiple local regions, facilitating subsequent processing. Among them, different local regions contain different object or scene features, and these features can be analyzed and processed more meticulously through slicing.

[0103] For example, assume that the deep spatio-temporal vector represents the depth information of a three-dimensional scene. Then, according to certain rules, the deep spatio-temporal vector is divided into multiple small blocks in the spatial dimension, and each small block is a depth vector block. Subsequently, the entire three-dimensional scene is divided into a grid, and the depth information corresponding to each grid cell constitutes a depth vector block. In this way, a complex three-dimensional scene is decomposed into multiple relatively simple local regions, and the features of each region can be analyzed and processed independently.

[0104] The process of concatenating multiple depth vector blocks in the sequence dimension is introduced below.

[0105] After obtaining multiple depth vector blocks by slicing in the spatial dimension, by concatenating them in the sequence dimension, the information of these local regions can be reorganized into an ordered sequence. This processing method can better preserve the temporal continuity and order of the depth video, and is also convenient for subsequent processing by the convolutional layer.

[0106] For example, for a depth video containing multiple frames, spatial slicing is performed on each frame to obtain multiple depth vector blocks, and then these depth vector blocks are arranged in chronological order to obtain a complete depth vector sequence. This depth vector sequence not only contains the depth information of each local region but also reflects the temporal changes of these regions.

[0107] The process of mapping the depth vector sequence block to the latent space is introduced below.

[0108] The convolutional layer is a commonly used layer structure in neural networks. The convolutional layer performs a sliding convolutional operation on the input data through a convolutional kernel to extract the local features of the data. In the above processing flow, the main role of the convolutional layer is to further process and transform the depth vector sequence, thereby mapping the depth vector sequence into the latent space.

[0109] Optionally, the convolutional layer performs a convolution operation on the sequence of depth vectors, extracting various features in the sequence through different convolutional kernels. These features can be local depth change patterns, shape features of objects, etc. During the convolution process, the convolutional layer also reduces the dimension and fuses the features. By setting appropriate numbers of convolutional kernels and strides, the dimension of the features can be reduced while different local features are fused together to form a more abstract and compact feature representation. The above process can remove some redundant information, improving the quality and expressive ability of the features. After being processed by the convolutional layer, the sequence of depth vectors is converted into a new feature representation, which is called the first long sequence representation for convenience of description. This first long sequence representation is located in the latent space and is an abstract representation of the depth video, containing the key features and information of the depth video in both time and space.

[0110] First, the VAE realizes spatio-temporal compression, significantly reducing the amount of data, lowering the computational cost, and improving the processing efficiency, enabling the model to handle large-scale data. Second, the operations of segmentation, splicing, and convolution effectively extract the local and global features of the depth video, and the generated first long sequence representation accurately reflects the essence of the depth video, facilitating the subsequent generation of a three-dimensional video with strong realism. Third, the dimension-by-dimension processing conforms to the spatio-temporal characteristics of the depth video and can comprehensively grasp the structure of the depth video. Moreover, the VAE encodes based on probability distributions, enhancing the generalization ability of the model and enabling it to adapt to diverse depth scenarios. Finally, this long sequence representation can be fused with the information of other modules to collaboratively guide video generation, making the output content richer, more accurate, meeting expectations, and improving the video generation efficiency.

[0111] In step S303, through the second signal processing module in the video generation model, the reference image among multiple data signals is processed to obtain a second long sequence representation, which is the long sequence representation of the reference image.

[0112] In the embodiments of the present disclosure, the reference image is used to provide important visual information and a reference basis for video generation. Optionally, the reference image can be a static image or a group of images with specific associations. The characteristic of the reference image is that it contains information such as specific scenes, object appearances, and color styles required for video generation.

[0113] For example, when generating a video with a specific architectural style, the reference image can be a photo of the exterior of the building, and the model can learn features such as the shape, color, and texture of the building from the reference image so as to present a similar visual effect when generating the video.

[0114] In some embodiments, the second signal processing module includes a variational autoencoder. The reference image is compressed by the variational autoencoder, then segmented and spliced, and finally mapped to the latent space through a convolutional layer to obtain the second long sequence representation. Correspondingly, first, for the reference image among multiple data signals, the reference image is input into the second signal processing module in the video generation model, and the reference image includes object information. Secondly, the reference image is spatially compressed by the variational autoencoder in the second signal processing module to obtain an image spatial vector. Thirdly, through the second signal processing module, the image spatial vector is segmented in the spatial dimension to obtain multiple image vector blocks. Fourthly, through the second signal processing module, multiple image vector blocks are spliced in the sequence dimension to obtain an image vector sequence. Finally, through the convolutional layer in the second signal processing module, the image vector sequence block is mapped to the latent space to obtain the second long sequence representation, and the second long sequence representation is the long sequence representation of the reference image.

[0115] The following introduces the process of spatially compressing the reference image.

[0116] The variational autoencoder consists of an encoder and a decoder. The encoder can map the input reference image to a latent space, thereby achieving spatial compression of the reference image. The encoder processes the reference image through a series of neural network layers (such as convolutional layers) to learn the feature distribution of the image. The encoder can convert the high-dimensional pixel information of the image into a low-dimensional latent variable representation, that is, an image spatial vector. In this process, the VAE can capture the key features of the image while removing redundant information, thereby achieving spatial compression.

[0117] The following introduces the process of spatially compressing the reference image.

[0118] After obtaining the image spatial vector, the second signal processing module will segment it in the spatial dimension to obtain multiple image vector blocks. By segmenting the image spatial vector, the overall image features can be decomposed into multiple local features, which is convenient for subsequent more detailed processing and analysis. Different image vector blocks may correspond to different regions in the reference image or different parts of an object, and can better capture the local details and features of the image. Optionally, the segmentation method can be designed according to specific requirements. For example, it can be divided according to a regular grid, and the image spatial vector can be evenly divided into multiple small blocks; it can also be divided according to the semantic information of the image, and the features belonging to the same object or region are divided into the same image vector block.

[0119] The following introduces the process of splicing multiple image vector blocks in the sequence dimension.

[0120] After completing the segmentation of the spatial dimension, multiple image vector blocks are then spliced ​​in the sequence dimension to obtain an image vector sequence. By splicing the segmented image vector blocks in the sequence dimension, these local features can be organized in a certain order to form an ordered sequence. The above sequence structure can better preserve the spatial information of the image and the association between the features, and it is also convenient for the subsequent convolution layer processing, because the convolution layer can effectively extract and transform the sequence data. Optionally, the splicing order can be determined according to the spatial position relationship of the image space vector, for example, it can be spliced ​​in the order from left to right and from top to bottom to ensure that adjacent image vector blocks in the sequence are also adjacent in the original image.

[0121] The following describes the process of mapping image vector sequence blocks into latent space.

[0122] After being processed by the convolutional layer, the image vector sequence is converted into a new feature representation, which is called the second long sequence representation for ease of explanation. The second long sequence representation is located in the latent space and is an abstract representation of the reference image. The second long sequence representation retains the key features and information of the reference image and exists in a more compact form that is more suitable for model processing. The second long sequence representation can be used as an important input for the subsequent video generation process, helping the model to refer to the features and style of the reference image when generating videos and generate video content that is more in line with expectations.

[0123] The reference image is spatially compressed by the variational autoencoder, which reduces the amount of image data and makes the model operation more efficient, so that a large number of reference images can be processed quickly. The image space vector is disassembled by the segmentation operation according to the spatial dimension, so that the model can focus on the object information in each area and mine local details; the image information can be ensured to be coherent by integrating it according to the sequence dimension through the splicing operation. The image vector sequence is further refined through the convolution layer, and the image vector sequence is converted into the second longest sequence representation to condense the key features of the reference image. The above method improves the accuracy and richness of the video content, makes the generated results more in line with expectations, and improves the efficiency of video generation.

[0124] In step S304, the camera trajectory video in the multiple data signals is processed by the third signal processing module in the video generation model to obtain a third long sequence representation, where the third long sequence representation is a long sequence representation of the camera trajectory video.

[0125] In the disclosed embodiment, the camera trajectory video records the movement information of the camera during the shooting process. The camera trajectory video contains information such as the position of the camera at different times (such as coordinates in three-dimensional space), orientation (pitch angle, yaw angle, roll angle, etc.), and possible movement speed and acceleration. The above information is crucial for generating videos with real perspective changes and a sense of movement, because the movement of the camera directly affects the scene content and the presentation of the picture seen by the audience.

[0126] In some embodiments, the third signal processing module obtains a camera trajectory vector through Plücker embedding, and then splits and splices it, and finally maps it to the latent space through a convolution layer to obtain a third long sequence representation. Correspondingly, first, for the camera trajectory video in the multiple data signals, the camera trajectory video is input into the third signal processing module in the video generation model, and the camera trajectory video includes camera trajectory information. Secondly, through the third signal processing module, the camera trajectory video is subjected to Plücker embedding to obtain a camera trajectory vector. Thirdly, through the third signal processing module, the camera trajectory vector is split in the spatial dimension to obtain a plurality of camera trajectory vector blocks. Thirdly, through the third signal processing module, a plurality of camera trajectory vector blocks are spliced ​​in the sequence dimension to obtain a camera trajectory vector sequence. Finally, through the convolution layer in the third signal processing module, the camera trajectory vector sequence block is mapped to the latent space to obtain a third long sequence representation, and the third long sequence representation is a long sequence representation of the camera trajectory video.

[0127] The following describes the process of Plücker embedding for camera trajectory videos.

[0128] Plücker embedding is a mathematical method that maps geometric objects (mainly the motion trajectory of the camera in camera trajectory processing) to a high-dimensional vector space. In three-dimensional space, the motion trajectory of the camera can be regarded as a series of straight lines or curve segments. Plücker embedding can convert these geometric information into vector representations, namely camera trajectory vectors. Among them, Plücker coordinates are a coordinate system used to represent straight lines in three-dimensional space. The camera's motion trajectory is decomposed into a series of straight line segments, and these straight line segments are represented by Plücker coordinates, and then converted into vector form through specific embedding operations. The advantage of the above method is that it converts complex geometric trajectory information into vector data that is easy for computers to process, which is convenient for subsequent numerical calculations and feature extraction. Plücker embedding can effectively capture the geometric features of the camera trajectory, including information such as the direction and curvature of the trajectory, and can retain these features in high-dimensional space, providing richer and more accurate information for subsequent processing.

[0129] The following describes the process of segmenting the camera trajectory vector in the spatial dimension.

[0130] After obtaining the camera trajectory vector, the camera trajectory vector can be sliced in the spatial dimension to obtain multiple camera trajectory vector blocks. Through slicing, the overall camera trajectory information can be decomposed into multiple local information, which is convenient for more detailed analysis and processing. Different camera trajectory vector blocks correspond to the trajectory information of the camera in different spatial regions or different motion stages, so that the local characteristics and changes of the trajectory can be better captured. Optionally, the slicing method can be designed according to specific requirements. For example, the camera trajectory vector can be evenly sliced at a fixed time interval or spatial distance; it can also be unevenly sliced according to the characteristic points of the camera motion (such as acceleration points, turning points, etc.), and the parts with similar motion characteristics are divided into one camera trajectory vector block.

[0131] The process of splicing multiple camera trajectory vector blocks in the sequence dimension is introduced below.

[0132] After completing the slicing in the spatial dimension, multiple camera trajectory vector blocks can be spliced in the sequence dimension to obtain a camera trajectory vector sequence. Through splicing, these local camera trajectory information can be reorganized in chronological order to form an ordered sequence. The above sequence structure can better retain the time continuity and order of the camera trajectory, and is also convenient for the subsequent processing of the convolutional layer, because the convolutional layer can effectively extract and transform the sequence data. Optionally, the splicing order can be in the chronological order of the camera trajectory vector blocks in the original camera trajectory, which can ensure that adjacent camera trajectory vector blocks in the sequence are also adjacent in time, so as to accurately reflect the motion process of the camera.

[0133] The process of mapping the camera trajectory vector sequence block to the latent space is introduced below.

[0134] After being processed by the convolutional layer, the camera trajectory vector sequence is converted into a new feature representation, which is called the third long sequence representation for convenience of description. This third long sequence representation is located in the latent space and is an abstract representation of the camera trajectory video. This third long sequence representation retains the key features and information of the camera trajectory, and at the same time exists in a more compact and more suitable form for model processing. The third long sequence representation can be an important input for the subsequent video generation process, helping the model to generate videos with real perspective changes and motion effects according to the motion trajectory of the camera.

[0135] The complex camera trajectory information is converted into vectors through Plücker embedding, effectively simplifying the data form, facilitating subsequent processing, and improving the computational efficiency. By splitting the camera trajectory vectors in the spatial dimension, the camera motion details at different spatial positions can be focused on to capture local features. By concatenating the vectors in the sequence dimension, the temporal coherence of the camera motion is ensured, and the entire trajectory is presented completely. The convolutional layer further mines the feature patterns in the camera trajectory vector sequence, maps them to the latent space, generates the third-longest sequence representation, and accurately extracts the key information of the camera trajectory. The third-longest sequence representation can provide an accurate camera motion reference for video generation, making the perspective switching and scene movement in the generated video more natural and smooth, highly conforming to the actual shooting effect, greatly enhancing the realism and visual experience of the video, and improving the video generation efficiency.

[0136] In step S305, the text encoder in the video generation model extracts features from the text prompt signal to obtain text features.

[0137] In the embodiments of the present disclosure, the text encoder is an important component in the video generation model for processing text data. The main role of the text encoder is to convert the input natural language text into a numerical feature representation that can be understood and processed by a computer. The above feature representation can capture information such as semantics, grammar, and context in the text, providing necessary guidance and constraints for the subsequent video generation process. The text encoder is usually constructed based on deep learning models, and common ones include those based on recurrent neural networks (RNNs) and their variants (such as long short-term memory networks LSTMs, gated recurrent units GRUs), convolutional neural networks (CNNs), and encoders of the Transformer architecture. The embodiments of the present disclosure do not limit the text encoder.

[0138] The text features are used to provide text prompt information. The process of extracting features from the text prompt signal is introduced below.

[0139] First, tokenize the text prompt signal, splitting the input text prompt signal into individual words or sub-word units. Then, assign a low-dimensional vector representation, i.e., word embedding, to each tokenized word or sub-word. Next, input the text sequence after word embedding into the text encoder. Encoders with different architectures will process the input sequence in different ways. Inside the encoder, a series of transformation and aggregation operations are performed on the input word vectors. For example, the multi-head self-attention layer in the Transformer architecture calculates the attention weights between each word vector and other word vectors, and then performs a weighted sum of the word vectors according to these weights to obtain a context-aware representation for each word. After multiple layers of transformation and aggregation, the encoder outputs a feature representation of the text. To obtain the feature representation of the entire text prompt signal, the features output by the encoder are usually further processed. A common method is to take the vector at the first position of the output of the last layer of the encoder (in the Transformer architecture, usually the vector corresponding to the [CLS] token) as the global feature representation of the text, and this vector synthesizes the information of all words in the text.

[0140] In step S306, the video diffusion model in the video generation model is used to process multiple long sequence representations and text features to obtain the target video features.

[0141] In the embodiments of the present disclosure, the multiple long sequence representations come from different signal processing modules, such as long sequence representations obtained after processing depth videos, reference images, camera trajectory videos, etc. The multiple long sequence features contain multi-dimensional spatio-temporal information such as the depth of the video, the appearance of the image, and the camera movement, and are abstract expressions of different levels of the original data. The text input by the user describing the content, style, theme, etc. of the target video is called the text prompt signal. By extracting features from the text prompt signal, the text features are obtained.

[0142] The video diffusion model first encodes multiple long sequence representations and converts them into the form of feature vectors that can be processed by the model. Then, using techniques such as the attention mechanism, these encoded features are fused, weights are assigned to each feature according to different task requirements, and their importance in generating the target video features is determined. It should be noted that before inputting into the video diffusion model, the above-mentioned multiple long sequence representations can be aligned. Optionally, the multiple long sequence representations are aligned by adjusting the convolutional layer. The embodiments of the present disclosure do not limit this. Then, based on the fused features, the video diffusion model performs a diffusion operation in the latent space, gradually adding noise to the data to construct a Markov chain from real data to pure noise. In the reverse process, the video diffusion model learns to denoise, and based on the fused features as conditions, gradually recovers the target video features from the noise. Each step of denoising adjusts the feature generation direction according to the current conditions. Finally, after multiple rounds of iterative denoising, the video diffusion model outputs the target video features. The target video features integrate the multi-dimensional information of the long sequence representations and the semantic information of the text prompts, and contain feature descriptions of various aspects such as the content, structure, and motion of the target video, providing a key basis for generating the complete target video subsequently.

[0143] In some embodiments, the steps of processing multiple long sequence representations by the video diffusion model in the video generation model to obtain the target video features include: splicing multiple long sequence representations in the sequence direction to obtain a multi-modal fusion feature; inputting the multi-modal fusion feature into the video diffusion model; injecting text features into the video diffusion model through the cross-attention mechanism; and processing the multi-modal fusion feature and the text features based on the full attention mechanism by the video diffusion model in the video generation model to obtain the target video features.

[0144] Among them, by splicing multiple long sequence representations in the sequence direction, information of different modalities can be integrated together. In this way, the originally scattered features are combined into a continuous feature sequence containing multi-faceted information, forming a multi-modal fusion feature.

[0145] The cross-attention mechanism allows the model to focus on the key parts of the text features when processing the multi-modal fusion features. The cross-attention mechanism calculates the correlation between the multi-modal fusion features and the text features, assigns a weight to each element in the text features, and then integrates the text features into the processing of the multi-modal fusion features according to these weights. Injecting text features into the video diffusion model through the cross-attention mechanism enables the video diffusion model to make targeted adjustments according to the text prompts when generating video features, making the generated video more in line with the user's expectations.

[0146] The full attention mechanism enables the model to comprehensively consider the mutual relationships among various elements in the feature sequence when processing multi-modal fusion features and text features. The full attention mechanism can not only capture the dependencies among elements within the same modality but also capture the interaction information among elements of different modalities. During the reverse denoising process of the video diffusion model, the full attention mechanism calculates the denoising direction and intensity for each time step based on the current noise state and the input multi-modal fusion features and text features. Through continuous iterative updates, the model gradually recovers the target video features from the noise. This feature integrates various information contained in the multi-modal fusion features and the semantic requirements conveyed by the text features, and is a complete and accurate representation of the target video features.

[0147] Among them, the order of concatenating multiple long sequence representations can be: text features + long sequence representation of the camera trajectory video + long sequence representation of the reference image + long sequence representation of the depth video + long sequence representation of the noise video. It can also be other concatenation orders, and the embodiments of the present disclosure do not limit this.

[0148] By concatenating multiple long sequence representations in the sequence direction, multi-source information such as depth, image, and trajectory is integrated, providing a comprehensive basis for video generation and enriching the video content dimension. By inputting the multi-modal fusion features into the video diffusion model, the model can perform diffusion and denoising based on multiple information sources to generate video features that better fit complex scenarios. By injecting text features into the video diffusion model through the cross-attention mechanism, the generation direction is accurately guided to ensure that the video meets the user's semantic requirements. By processing the multi-modal fusion features and text features based on the full attention mechanism through the video diffusion model, the feature interaction is comprehensively considered. It not only captures the associations among elements within each modality but also explores the connections between different modalities, enabling the elements of the video to cooperate with each other to generate target video features with logical coherence and rich details, significantly improving the quality and accuracy of the generated video and enhancing the video generation efficiency.

[0149] Next, the structure of the video diffusion model will be introduced. Refer to Figure 4 as shown Figure 4 is a schematic structural diagram of a video diffusion model provided according to an exemplary embodiment. As Figure 4 shown, the noise features and text features are input into the first basic block. Then, the output of the first basic block is input into the second basic block. The second basic block includes a two-dimensional self-attention layer, a three-dimensional self-attention layer, a cross-attention layer, and a feedforward neural network (FFN). The output of the second basic block is used as the input of the third basic block. The output of the third basic block is the target video features.

[0150] Note that when a certain control condition is not input, that is, when a certain data signal is empty, the data signal can be set to 0. It is also possible to directly not input this control condition, as long as the position encoding used for each control condition is consistent with that during model training. If extended to other control conditions, a new long sequence representation of the control condition can be directly connected after the multi-modal fusion features.

[0151] In step S307, the target video feature is decoded to obtain the target video.

[0152] In the embodiments of the present disclosure, this step is the same as step S204 above. Refer to step S204 above, and the present disclosure will not repeat it here.

[0153] The embodiments of the present disclosure provide a video generation solution. Multiple signal processing modules in the video generation model process multiple data signals to generate multiple long sequence representations. The above multiple long sequence representations cover at least one type of information such as noise, camera trajectory, depth, and objects. By integrating multiple control conditions in the sequence dimension, it not only provides a rich and key data source for the video diffusion model, enabling the video diffusion model to understand videos or images from multiple dimensions, but also helps to accurately capture video details and features through the comprehensive processing of multiple data signals, thereby generating high-quality target video features, and then obtaining the target video, improving the quality of the generated video and the video generation efficiency.

[0154] Figure 5 is a block diagram of a video generation device shown according to an exemplary embodiment. As Figure 5 shown, the device includes: a first processing unit 501, a second processing unit 502, and a decoding unit 503.

[0155] The first processing unit 501 is configured to process multiple data signals respectively through multiple signal processing modules in the video generation model to obtain multiple long sequence representations. The multiple signal processing modules are used to process the input video or image into corresponding long sequence representations, and each long sequence representation corresponds to one data signal. The multiple data signals are used to provide at least one of noise information, camera trajectory information, depth information, and object information;

[0156] The second processing unit 502 is configured to process the multiple long sequence representations through the video diffusion model in the video generation model to obtain the target video feature. The video diffusion model is used to output video features based on the input data;

[0157] The decoding unit 503 is configured to decode the target video feature to obtain the target video.

[0158] In some embodiments, the first processing unit 501 is configured to, for the depth video among multiple data signals, input the depth video into the first signal processing module in the video generation model. The depth video includes depth information. Compress the depth video temporally and spatially through the variational autoencoder in the first signal processing module to obtain a depth spatio-temporal vector. Split the depth spatio-temporal vector in the spatial dimension through the first signal processing module to obtain multiple depth vector blocks. Concatenate the multiple depth vector blocks in the sequence dimension through the first signal processing module to obtain a depth vector sequence. Map the depth vector sequence block to the latent space through the convolutional layer in the first signal processing module to obtain a first long sequence representation, and the first long sequence representation is the long sequence representation of the depth video.

[0159] In some embodiments, the first processing unit 501 is configured to, for the reference image among multiple data signals, input the reference image into the second signal processing module in the video generation model. The reference image includes object information. Compress the reference image spatially through the variational autoencoder in the second signal processing module to obtain an image spatial vector. Split the image spatial vector in the spatial dimension through the second signal processing module to obtain multiple image vector blocks. Concatenate the multiple image vector blocks in the sequence dimension through the second signal processing module to obtain an image vector sequence. Map the image vector sequence block to the latent space through the convolutional layer in the second signal processing module to obtain a second long sequence representation, and the second long sequence representation is the long sequence representation of the reference image.

[0160] In some embodiments, the first processing unit 501 is configured to, for the camera trajectory video among multiple data signals, input the camera trajectory video into the third signal processing module in the video generation model. The camera trajectory video includes camera trajectory information. Perform Plücker embedding on the camera trajectory video through the third signal processing module to obtain a camera trajectory vector. Split the camera trajectory vector in the spatial dimension through the third signal processing module to obtain multiple camera trajectory vector blocks. Concatenate the multiple camera trajectory vector blocks in the sequence dimension through the third signal processing module to obtain a camera trajectory vector sequence. Map the camera trajectory vector sequence block to the latent space through the convolutional layer in the third signal processing module to obtain a third long sequence representation, and the third long sequence representation is the long sequence representation of the camera trajectory video.

[0161] In some embodiments, the second processing unit 502 is configured to concatenate multiple long sequence representations in the sequence direction to obtain a multi-modal fusion feature. Input the multi-modal fusion feature into the video diffusion model. Inject the text feature into the video diffusion model through the cross-attention mechanism. Process the multi-modal fusion feature and the text feature based on the full attention mechanism through the video diffusion model in the video generation model to obtain a target video feature.

[0162] In some embodiments, the apparatus further includes:

[0163] A third processing unit configured to extract features from the text prompt signal through a text encoder in the video generation model to obtain text features.

[0164] The embodiments of the present disclosure provide a video generation apparatus that processes various data signals through multiple signal processing modules in a video generation model to generate multiple long - sequence representations. The above - mentioned multiple long - sequence representations cover at least one type of information such as noise, camera trajectory, depth, and objects. By integrating various control conditions in the sequence dimension, it not only provides a rich and key data source for the video diffusion model, enabling the video diffusion model to understand videos or images from multiple dimensions, but also helps to accurately capture video details and features through the comprehensive processing of various data signals, thereby generating high - quality target video features, and then obtaining the target video, improving the quality of the generated video and the video generation efficiency.

[0165] It should be noted that the video generation apparatus provided in the above embodiments is only illustrated by the division of the above - mentioned functional units. In practical applications, the above functions can be assigned to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the video generation apparatus provided in the above embodiments and the embodiments of the video generation method belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be elaborated here.

[0166] Regarding the video generation apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0167] In the embodiments of the present disclosure, the electronic device can be a terminal or a server. When the electronic device is a terminal, the terminal is used as the execution subject to implement the technical solutions provided in the embodiments of the present disclosure; when the electronic device is a server, the server is used as the execution subject to implement the technical solutions provided in the embodiments of the present disclosure; or, the technical solutions provided in the present disclosure are implemented through the interaction between the terminal and the server. The embodiments of the present disclosure do not limit this.

[0168] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment. Generally, the electronic device 600 includes a processor 601 and a memory 602.

[0169] The processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0170] The memory 602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 602 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 is used to store at least one program code, and the at least one program code is used to be executed by the processor 601 to implement the video generation method provided in the method embodiments of the present disclosure.

[0171] In some embodiments, the electronic device 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.

[0172] The peripheral device interface 603 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0173] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 604 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), and this disclosure does not limit this.

[0174] The display screen 605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input to the processor 601 as control signals for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which is disposed on the front panel of the electronic device 600; in other embodiments, there may be at least two display screens 605, which are respectively disposed on different surfaces of the electronic device 600 or are in a foldable design; in still other embodiments, the display screen 605 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the electronic device 600. Even more, the display screen 605 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 605 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0175] The camera module 606 is used to capture images or videos. Optionally, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the electronic device, and the rear camera is disposed on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 606 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0176] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.

[0177] The power supply 608 is used to supply power to each component in the electronic device 600. The power supply 608 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 608 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0178] Those skilled in the art can understand that Figure 6 the structure shown in does not limit the electronic device 600, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0179] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as the memory 602 including instructions. The above instructions can be executed by the processor 601 of the electronic device 600 to complete the above video generation method. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0180] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above video generation method.

[0181] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0182] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video generation method, characterized in that: The method comprises: Processing the multiple data signals respectively by multiple signal processing modules in the video generation model to obtain multiple long sequence representations, wherein the multiple signal processing modules are used to process the input video or image into corresponding long sequence representations, each long sequence representation corresponds to a data signal, and the multiple data signals are used to provide at least one of noise information, camera trajectory information, depth information and object information; The plurality of long sequence representations are processed by a video diffusion model in the video generation model to obtain target video features, wherein the video diffusion model is used to output video features based on input data; The target video feature is decoded to obtain the target video.

2. The video generation method according to claim 1, characterized in that: The multiple signal processing modules in the video generation model process the multiple data signals respectively to obtain multiple long sequence representations, including: For a depth video among the multiple data signals, input the depth video into a first signal processing module in the video generation model, wherein the depth video includes the depth information; Compressing the depth video in time and space by a variational autoencoder in the first signal processing module to obtain a depth spatiotemporal vector; By means of the first signal processing module, the depth spatiotemporal vector is segmented in the spatial dimension to obtain a plurality of depth vector blocks; By means of the first signal processing module, the plurality of depth vector blocks are spliced ​​in a sequence dimension to obtain a depth vector sequence; The depth vector sequence block is mapped to a latent space through a convolutional layer in the first signal processing module to obtain a first long sequence representation, where the first long sequence representation is a long sequence representation of the depth video.

3. The video generation method according to claim 1, characterized in that: The multiple signal processing modules in the video generation model process the multiple data signals respectively to obtain multiple long sequence representations, including: For a reference image in the plurality of data signals, inputting the reference image into a second signal processing module in the video generation model, wherein the reference image includes the object information; Compressing the reference image spatially by a variational autoencoder in the second signal processing module to obtain an image space vector; By means of the second signal processing module, the image space vector is segmented in the space dimension to obtain a plurality of image vector blocks; splicing the plurality of image vector blocks in a sequence dimension through the second signal processing module to obtain an image vector sequence; The image vector sequence block is mapped to a latent space through a convolutional layer in the second signal processing module to obtain a second long sequence representation, where the second long sequence representation is a long sequence representation of the reference image.

4. The video generation method according to claim 1, characterized in that: The multiple signal processing modules in the video generation model process the multiple data signals respectively to obtain multiple long sequence representations, including: For a camera trajectory video among the multiple data signals, inputting the camera trajectory video into a third signal processing module in the video generation model, wherein the camera trajectory video includes the camera trajectory information; Performing Plücker embedding on the camera trajectory video through the third signal processing module to obtain a camera trajectory vector; By means of the third signal processing module, the camera trajectory vector is segmented in the spatial dimension to obtain a plurality of camera trajectory vector blocks; By means of the third signal processing module, the plurality of camera trajectory vector blocks are spliced ​​in a sequence dimension to obtain a camera trajectory vector sequence; The camera trajectory vector sequence block is mapped to a latent space through a convolutional layer in the third signal processing module to obtain a third long sequence representation, where the third long sequence representation is a long sequence representation of the camera trajectory video.

5. The video generation method according to claim 1, characterized in that: The step of processing the plurality of long sequence representations through the video diffusion model in the video generation model to obtain target video features includes: Splicing the multiple long sequence representations in the sequence direction to obtain multimodal fusion features; Inputting the multimodal fusion features into the video diffusion model; Injecting text features into the video diffusion model through a cross-attention mechanism, wherein the text features are used to provide text prompt information; The multimodal fusion features and the text features are processed based on the full attention mechanism through the video diffusion model in the video generation model to obtain the target video features.

6. The video generation method according to any one of claims 1 to 5, characterized in that: The method further comprises: The text feature is obtained by extracting features from the text prompt signal through the text encoder in the video generation model.

7. A video generating device, characterized in that: The device comprises: A first processing unit is configured to process the multiple data signals respectively through multiple signal processing modules in the video generation model to obtain multiple long sequence representations, wherein the multiple signal processing modules are used to process the input video or image into corresponding long sequence representations, each long sequence representation corresponds to a data signal, and the multiple data signals are used to provide at least one of noise information, camera trajectory information, depth information and object information; A second processing unit is configured to process the multiple long sequence representations through a video diffusion model in the video generation model to obtain target video features, wherein the video diffusion model is used to output video features based on input data; The decoding unit is configured to decode the target video feature to obtain the target video.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the video generating method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generating method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the video generating method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Video generation method and device, electronic equipment and storage medium

    CN120434476A