Video generation method, device, electronic device and storage medium

Generating videos through multimodal data solves the problem of low video generation quality, achieves higher flexibility and controllability, and improves the accuracy of the generated results.

CN120434476BActive Publication Date: 2025-09-30BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510933855.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-30
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The video generation quality in the existing technology is low and the controllability is poor, resulting in the generation results not meeting expectations.

Method used

Multimodal data is used to generate videos. By obtaining a video generation request, determining task information and visual description information, generating a contextual condition control sequence, and inputting the multimodal data into a video generation model, the target video is generated.

Benefits of technology

The flexibility and controllability of video generation are improved, and the accuracy and quality of generated results are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434476B_ABST
    Figure CN120434476B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, electronic device, and storage medium. The method comprises: obtaining a video generation request, wherein the video generation request includes multimodal data input by a user; determining, based on the multimodal data, task information of a video generation task corresponding to the video generation request and visual description information of the video generation task; generating a context condition control sequence based on the multimodal data, and generating a content feature sequence based on the context condition control sequence; and inputting the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video. The present disclosure can improve the quality of video generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video technology, and in particular to a video generation method, device, electronic device, and storage medium. Background Art

[0002] With the development of artificial intelligence technology, video generation has become an important research direction in the field of artificial intelligence and has received widespread attention in recent years.

[0003] Traditional technologies typically use a single-modal input approach to achieve video generation. For example, text content or image sequences are directly input into an AI model, which then generates video content based on the text or image sequence. However, using single-modal data to generate video results has poor controllability, often resulting in inconsistent results and low video quality. Summary of the Invention

[0004] The present disclosure provides a video generation method, device, electronic device, and storage medium to at least address the problem of low video quality in related technologies. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a video generation method is provided, including:

[0006] Obtaining a video generation request, where the video generation request includes multimodal data input by a user;

[0007] Determining, based on the multimodal data, task information of a video generation task corresponding to the video generation request and visual description information of the video generation task;

[0008] generating a context condition control sequence based on the multimodal data, and generating a content feature sequence based on the context condition control sequence;

[0009] The task information, the visual description information, and the content feature sequence are input into a video generation model to generate a target video.

[0010] In one embodiment, determining, based on the multimodal data, task information of a video generation task corresponding to the video generation request and visual description information of the video generation task includes:

[0011] Inputting the multimodal data into an intent recognition model to obtain task information of a video generation task corresponding to the video generation request;

[0012] The task information and the multimodal data are input into a content expansion model to obtain visual description information of the video generation task.

[0013] In one embodiment, the multimodal data includes base content data and data indicating a spatiotemporal position; and generating a contextual conditional control sequence based on the multimodal data includes:

[0014] Encoding the base content data and the data indicating the spatiotemporal position respectively to obtain a first coding sequence corresponding to the base content data and a second coding sequence corresponding to the data indicating the spatiotemporal position;

[0015] The first coding sequence and the second coding sequence are superimposed to obtain a context condition control sequence.

[0016] In one embodiment, superimposing the first coding sequence and the second coding sequence to obtain a context condition control sequence includes:

[0017] The codes at the same position in the first coding sequence and the second coding sequence are superimposed to obtain a context condition control sequence.

[0018] In one embodiment, the generating of the content feature sequence based on the context condition control sequence includes:

[0019] The context condition control sequence and the noise feature sequence are concatenated to obtain a content feature sequence.

[0020] In one embodiment, the multimodal data includes time series multimodal data and / or non-time series multimodal data.

[0021] In one embodiment, after inputting the task information, the visual description information, and the content feature sequence into a video generation model, the method further includes:

[0022] For a target sequence in the content feature sequence, determining a target index range based on attribute information corresponding to the target sequence; wherein the target sequence is a feature sequence corresponding to non-time series multimodal data;

[0023] Based on the target index range, index information is assigned to the target sequence.

[0024] In one embodiment, the video generation model includes multiple processing layers; after inputting the task information, the visual description information, and the content feature sequence into the video generation model, the method further includes:

[0025] For each processing layer of the video generation model, in the process of performing attention calculation on the noise feature sequence through the processing layer, a target conditional frame is selected from the stored conditional frames according to a preset conditional frame selection strategy;

[0026] Attention calculation is performed based on the feature sequence corresponding to the target condition frame and the noise feature sequence.

[0027] In one embodiment, inputting the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video includes:

[0028] Inputting the task information, the visual description information, and the content feature sequence into a video generation model;

[0029] superimposing the task feature sequence corresponding to the task information and the content feature sequence through the video generation model to obtain an updated content feature sequence;

[0030] A target video is generated by the video generation model based on the updated content feature sequence, the task information and the visual description information.

[0031] According to a second aspect of an embodiment of the present disclosure, there is provided a video generating apparatus, including:

[0032] an acquiring unit configured to execute acquiring a video generation request, wherein the video generation request includes multimodal data input by a user;

[0033] A first determining unit is configured to determine task information of a video generation task corresponding to the video generation request and visual description information of the video generation task based on the multimodal data;

[0034] A first generating unit is configured to generate a context condition control sequence based on the multimodal data, and generate a content feature sequence based on the context condition control sequence;

[0035] The second generating unit is configured to input the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video.

[0036] In one embodiment, the first determining unit includes:

[0037] A first determining subunit is configured to input the multimodal data into an intent recognition model to obtain task information of a video generation task corresponding to the video generation request;

[0038] The second determining subunit is configured to input the task information and the multimodal data into a content expansion model to obtain visual description information of the video generation task.

[0039] In one embodiment, the multimodal data includes base content data and data indicating a spatiotemporal position; and the first generating unit includes:

[0040] an encoding subunit configured to perform encoding processing on the base content data and the data indicating the spatiotemporal position, respectively, to obtain a first encoding sequence corresponding to the base content data and a second encoding sequence corresponding to the data indicating the spatiotemporal position;

[0041] The first superposition subunit is configured to superpose the first coding sequence and the second coding sequence to obtain a context condition control sequence.

[0042] In one embodiment, the first superposition subunit is configured to perform:

[0043] The codes at the same position in the first coding sequence and the second coding sequence are superimposed to obtain a context condition control sequence.

[0044] In one embodiment, the first generating unit includes:

[0045] The splicing subunit is configured to splice the context condition control sequence and the noise feature sequence to obtain a content feature sequence.

[0046] In one embodiment, the multimodal data includes time series multimodal data and / or non-time series multimodal data.

[0047] In one embodiment, the apparatus further comprises:

[0048] A second determining unit is configured to determine a target index range for a target sequence in the content feature sequence based on attribute information corresponding to the target sequence; wherein the target sequence is a feature sequence corresponding to non-time series multimodal data;

[0049] An allocating unit is configured to allocate index information to the target sequence based on the target index range.

[0050] In one embodiment, the apparatus further comprises:

[0051] a screening unit configured to execute, for each processing layer of the video generation model, a process of performing attention calculation on the noise feature sequence through the processing layer, and to screen a target conditional frame from the stored conditional frames according to a preset conditional frame screening strategy;

[0052] A computing unit is configured to perform attention calculation based on the feature sequence corresponding to the target condition frame and the noise feature sequence.

[0053] In one embodiment, the second generating unit includes:

[0054] an input subunit, configured to input the task information, the visual description information, and the content feature sequence into a video generation model;

[0055] A second superposition subunit is configured to execute, through the video generation model, superimposing the task feature sequence corresponding to the task information and the content feature sequence to obtain an updated content feature sequence;

[0056] The generation subunit is configured to generate a target video based on the updated content feature sequence, the task information and the visual description information through the video generation model.

[0057] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0058] processor;

[0059] a memory for storing instructions executable by the processor;

[0060] The processor is configured to execute the instructions to implement the video generation method as described in any one of the first aspects.

[0061] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generation method as described in any one of the first aspects.

[0062] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes instructions. When the instructions are executed by a processor of an electronic device, the electronic device can execute the video generation method as described in any one of the first aspects.

[0063] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects: obtaining a video generation request, the video generation request includes multimodal data input by the user, determining the task information of the video generation task corresponding to the video generation request and the visual description information of the video generation task based on the multimodal data, generating a context condition control sequence based on the multimodal data, and generating a content feature sequence based on the context condition control sequence, and then inputting the task information, visual description information, and content feature sequence into the video generation model to generate the target video. Through the above scheme, it is possible to generate video content based on multimodal data, and through the flexible combination of multiple modal data, the flexibility of video generation can be improved. In addition, the multimodal data can more accurately reflect the needs of video generation, and the task information, visual description information, and the content feature sequence obtained by the context condition control sequence can form effective conditional control in the video generation process, thereby improving the controllability of the generated video content and the accuracy of the generation results, thereby improving the quality of video generation.

[0064] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0066] Figure 1 The figure is a flowchart of a video generating method according to an exemplary embodiment.

[0067] Figure 2 The figure is a flowchart showing a method of generating task information and visual description information according to an exemplary embodiment.

[0068] Figure 3 The figure is a schematic diagram showing generation of a content feature sequence according to an exemplary embodiment.

[0069] Figure 4 The figure is a flowchart showing a method of generating a target video according to an exemplary embodiment.

[0070] Figure 5 It is a schematic diagram of a framework of a video generation method according to an exemplary embodiment.

[0071] Figure 6 The figure is a block diagram of a video generating apparatus according to an exemplary embodiment.

[0072] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0073] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0074] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0075] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0076] Figure 1 This is a flow chart of a video generation method according to an exemplary embodiment. This embodiment uses the method applied to a terminal as an example. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. It is understandable that the method can be applied to any terminal with data processing capabilities, and this embodiment does not limit it. Figure 1 As shown, the specific steps include the following steps.

[0077] In step S110 , a video generation request is obtained.

[0078] The video generation request includes multimodal data input by the user.

[0079] In an embodiment of the present disclosure, the terminal may receive a video generation request input by a user through an input device, and the video generation request may include multimodal data input by the user. The multimodal data may include a variety of different types of data, for example, including but not limited to instruction data, image data, video data, and voice data. The instruction data may be used to describe the expectations or requirements for video generation. The specific content of the multimodal data may be input by the user according to habits and actual needs. The user may flexibly combine various types of data and freely implement various multimodal control tasks (such as video generation) through instruction control. In one example, the multimodal data may include video A, video B, and instruction data. The instruction data may be used to generate a new video based on the content of video A and the style of video B.

[0080] In step S120 , task information of the video generation task corresponding to the video generation request and visual description information of the video generation task are determined based on the multimodal data.

[0081] In an embodiment of the present disclosure, the terminal can perform intent analysis on the multimodal data to obtain task information of the video generation task corresponding to the video generation request. The task information may include a task identifier, such as a task identifier for a multi-image reference task, a task identifier for a video editing task, a task identifier for a video expansion task, a task identifier for a layout control task, etc. The task information may also include description information of the task, such as video A is a base video and video B is a reference style video. The terminal can also expand the prompt words based on the content of the multimodal data to obtain visual description information of the video generation task. The visual description information can be used to describe the rules or detailed content of the video generation.

[0082] In step S130 , a context condition control sequence is generated based on the multimodal data, and a content feature sequence is generated based on the context condition control sequence.

[0083] In an embodiment of the present disclosure, for each multimodal data item, the terminal may determine an encoder that matches the data type based on the data type of the multimodal data item, then input the multimodal data item into the encoder to obtain a coding sequence (i.e., a token sequence) for the multimodal data item, and then obtain a contextual condition control sequence based on the coding sequence. Based on the contextual condition control sequence and a preset sequence generation strategy, the terminal may generate a content feature sequence that includes the contextual condition control sequence.

[0084] In step S140 , the task information, visual description information, and content feature sequence are input into a video generation model to generate a target video.

[0085] In the disclosed embodiment, the terminal can input task information, visual description information, and content feature sequence into a video generation model to output a target video. Among them, the video generation model can be a DiT (Diffusion Transformers) model. The DiT model is a generation model that combines the diffusion model and the Transformer structure. The diffusion model is a probabilistic model that generates data through gradual denoising, while the Transformer structure has a powerful sequence modeling capability. By combining the two, DiT can achieve higher quality and flexibility in generation tasks. DiT can use contextual conditions and attention mechanisms to process input multimodal data, thereby providing better performance in tasks such as image and video generation. It can be understood that the video generation model in this embodiment can also adopt other base models with video generation diffusion functions, which is not limited in this embodiment.

[0086] In the above scheme, task information and visual description information can be generated based on data of multiple modalities, and context condition control sequences can also be generated based on data of multiple modalities. Then, the task information, visual description information, and content feature sequences obtained by the context condition control sequences are input into the video generation model for video generation. In this way, a flexible combination of multiple modal data can be achieved, which improves the flexibility of video generation. Moreover, multimodal data can more accurately reflect the needs of video generation, and task information, visual description information, and content feature sequences obtained by the context condition control sequences can form effective conditional control in the video generation process, thereby improving the controllability of generated video content and the accuracy of generation results, thereby improving the quality of video generation.

[0087] In an exemplary embodiment, Figure 2 As shown, based on the multimodal data, determining the task information of the video generation task corresponding to the video generation request and the visual description information of the video generation task includes:

[0088] In step S210 , the multimodal data is input into the intent recognition model to obtain task information of the video generation task corresponding to the video generation request.

[0089] In the embodiment of the present disclosure, the intent recognition model may be a Multimodal Large Language Model (MLLM). The terminal may input multimodal data into the intent recognition model and output task information of a video generation task corresponding to the video generation request.

[0090] In step S220 , the task information and multimodal data are input into the content expansion model to obtain visual description information of the video generation task.

[0091] In the disclosed embodiment, the content expansion model may be a PE (Position Embedding) model. The terminal may input task information and multimodal data into the content expansion model. The content expansion model may expand the prompt words based on the input task information and multimodal data, and then output text data containing multiple prompt words. Different output text paradigms may be used for different video generation tasks. The content expansion model may determine the corresponding text paradigm based on the task information, and then generate text data based on the text paradigm and the expanded prompt words. Optionally, the content expansion model may also be implemented through pre-trained text or other MLLM models, which is not limited in this embodiment.

[0092] In one example, a user inputs a picture containing a person and a picture containing a dog, and the instruction is to input a dog walking video. The content expansion model can expand the content based on the picture content and obtain the following content: a middle-aged man, wearing black clothes, walking on the road with a white puppy.

[0093] In the above scheme, user input data of multiple different modalities can be combined and uniformly transcribed into text form and input into the video generation model, realizing unified processing of multi-modal and multi-task, which can flexibly adapt to different task requirements and application scenarios, improve task compatibility, eliminate the need to train models for each task separately, and reduce parameter redundancy.

[0094] In an exemplary embodiment, multimodal data includes base content data and indicative data of spatiotemporal positions; generating a context condition control sequence based on the multimodal data includes: encoding the base content data and the indicative data of spatiotemporal positions respectively to obtain a first encoding sequence corresponding to the base content data and a second encoding sequence corresponding to the indicative data of spatiotemporal positions; and superimposing the first encoding sequence and the second encoding sequence to obtain a context condition control sequence.

[0095] In the disclosed embodiments, multimodal data may include multiple types of data, wherein the multimodal data may be divided into multimodal data of a time series type and / or multimodal data of a non-time series type according to whether the data has time series characteristics. The time series type indicates that the multimodal data has time series characteristics, and the characteristics need to be time-aligned during the calculation process. For example, the multimodal data of the time series type may be video, action data, camera movement data, etc. The multimodal data of the non-time series type indicates that the multimodal data does not have time series characteristics, and the characteristics do not need to be time-aligned during the calculation process. For example, the multimodal data of the non-time series type may be pictures, style information, lighting information, heavy camera movement, etc.

[0096] Furthermore, multimodal data of time series type can be used as constraints for model calculation. Therefore, according to the different constraint dimensions, it can be further divided into the following types: indicative data of spatiotemporal position, base content data, and data of other time series types. Among them, the indicative data of spatiotemporal position is data used to indicate the spatial position and / or temporal position of the target content. For example, the indicative data of spatiotemporal position can be the position of a person in the lower left corner of the video. The indicative data of spatiotemporal position can be specifically represented as a mask and / or a bounding box. The base content data can be the basic content for generating a video. For example, the instruction for video generation is: generate a new video based on the content of video A and the style of video B, then video A is the base content data. For another example, for a video expansion task, if the first 10 frames of video are given, then the first 10 frames are the base content data. Other time series types of data are other time series types of data other than the indicative data of spatiotemporal position and base content data, such as action data, camera trajectories, line drawings, etc.

[0097] After acquiring the multimodal data, the terminal identifies the type of each data in the multimodal data. If there is base content data and indication data of the spatiotemporal position, the base content data and the indication data of the spatiotemporal position can be encoded and processed respectively to obtain a first coding sequence corresponding to the base content data and a second coding sequence corresponding to the indication data of the spatiotemporal position. Then, the first coding sequence is superimposed on the second coding sequence to obtain a contextual condition control sequence. For different video generation tasks, the corresponding spatiotemporal position indication data (such as a mask) is also different. The terminal can determine the spatiotemporal position indication data corresponding to the task information of the video generation task based on the task information of the video generation task. Then, the determined spatiotemporal position indication data is input into the corresponding encoder to obtain a first coding sequence. It can be understood that in this embodiment, the coding sequence and the contextual condition control sequence are both token sequences.

[0098] Optionally, when encoding multimodal data, different types of encoders can be selected based on different data attributes. For example, a VAE (Variational Autoencoder) can be used for video data, and a text encoder can be used for text data. For other time-series or non-time-series data, a tokenizer encoder can be used.

[0099] Optionally, the terminal can identify the type of each multimodal data in a variety of ways. For example, the task information output by the intent recognition model can include description information of the task, such as video A is a base video and video B is a reference style video. This description information can reflect the type of each multimodal data. The terminal can identify the type of each multimodal data based on the task information. For another example, when the user inputs multimodal data, the user can configure the type of each multimodal data, and the terminal can set a type label for each multimodal data. In this way, in the subsequent processing process, the type of each multimodal data can be identified based on the type label.

[0100] In the above solution, by superimposing the first coding sequence and the second coding sequence, the spatiotemporal position of the target object in the base video can be controlled, thereby improving the controllability of video generation and thus improving the quality of video generation.

[0101] In an exemplary embodiment, superimposing the first coding sequence and the second coding sequence to obtain a context condition control sequence includes: superimposing the codes at the same position in the first coding sequence and the second coding sequence to obtain the context condition control sequence.

[0102] In the embodiment of the present disclosure, the terminal may perform positional superposition of codes at the same position in the first coding sequence and the second coding sequence to obtain a context condition control sequence.

[0103] In the above scheme, by superimposing tokens at the same position, the superposition of the first coding sequence and the second coding sequence can be achieved, thereby controlling the spatiotemporal position of the target object in the base video, improving the controllability of video generation, and thus improving the quality of video generation.

[0104] In an exemplary embodiment, generating a content feature sequence based on a context condition control sequence includes: concatenating the context condition control sequence and a noise feature sequence to obtain a content feature sequence.

[0105] In an embodiment of the present disclosure, the terminal can generate a content feature sequence through a content generation module. The content generation module may include a noise module, a context condition module and a DiT model. Among them, the noise module may be a noise video latent module, which is used to generate a noise feature sequence. The context condition module may include a timing condition submodule and a non-timing condition submodule. The timing condition submodule is used to encode the multimodal data of the timing type in the multimodal data, and the non-timing condition submodule is used to encode the multimodal data of the non-timing type. It can be understood that when there is at least one multimodal data, the context condition module can obtain at least one encoding sequence, and each encoding sequence is a context condition sequence. The content generation module can splice the noise feature sequence and all the context condition sequences to obtain a long token sequence, that is, a content feature sequence. It can be understood that the content feature sequence is a token sequence. Reference Figure 3 , taking the case where the spatiotemporal position indication data (i.e., mask and / or bounding box), base content data, other time series type data, and non-time series type data all exist in the multimodal data as an example, this embodiment provides a schematic diagram of content feature sequence generation. Among them, the noise feature sequence and each time series type multimodal data can use the same position coding. Specifically, the length of the token sequence of the noise feature sequence can be the same as the length of the token sequence of each time series type multimodal data, for example, both are 0~N. In this way, when the position coding is subsequently injected, the time dimension position coding of both can be set to 0~N. This method can enhance the timing alignment conditions and the frame-by-frame alignment capability of the generated video, optimize the performance in timing alignment, and thus improve the accuracy of video generation.

[0106] In this solution, the content generation module uniformly processes the various multimodal input data, unifying the multimodal control video generation task within a single framework. This eliminates the need for configuring additional models and reduces parameter redundancy. Furthermore, the DiT model can adopt a contextual conditional control paradigm. By tokenizing different control conditions and combining them with noise feature sequences into long token sequences, this solution can flexibly utilize different contextual conditions to implement various tasks, providing task flexibility and adaptability.

[0107] In an exemplary embodiment, after the task information, visual description information, and content feature sequence are input into the video generation model, the method further includes: determining a target index range for a target sequence in the content feature sequence based on attribute information corresponding to the target sequence; wherein the target sequence is a feature sequence corresponding to non-time series multimodal data; and assigning index information to the target sequence based on the target index range.

[0108] In an embodiment of the present disclosure, after the task information, visual description information, and content feature sequence are input into the video generation model, the video generation model can receive the task information output by the intent recognition model and determine the attribute information of the non-time series multimodal data based on the task information, that is, the attribute information of the feature sequence corresponding to the non-time series multimodal data. Among them, the attribute information can characterize the characteristics of the non-time series multimodal data, such as ID (identity information), style, lighting, re-camera, etc. Among them, ID can be used to characterize pictures, such as pictures of people, dogs, vehicles, etc.

[0109] Based on the attribute information, the video generation module can determine a target index range corresponding to the attribute information, and then assign index information to the target sequence based on the index within the target index range. Specifically, a global index can be assigned to the target sequence, or an index can be assigned to each token in the target sequence, which is not limited in this embodiment.

[0110] In the above scheme, assigning indexes within different index ranges to different attribute information can enable the DiT model to effectively identify the attributes corresponding to the target sequence in the subsequent calculation process. Since the attributes can reflect the characteristics of the task, the task recognition ability of the model is enhanced, enabling the model to recognize and combine different conditional attributes to achieve diversified tasks.

[0111] In an exemplary embodiment, the video generation model includes multiple processing layers; after the task information, visual description information, and content feature sequence are input into the video generation model, the method also includes: for each processing layer of the video generation model, in the process of performing attention calculation on the noise feature sequence through the processing layer, according to a preset conditional frame screening strategy, the target conditional frame is screened from the stored conditional frames; and attention calculation is performed based on the feature sequence corresponding to the target conditional frame and the noise feature sequence.

[0112] Among them, the conditional frame refers to the additional input information used to guide the content generation process in the attention calculation (such as videos, images, etc. in multimodal data input by users), which is usually provided in the form of frames (such as images, video clips or time series data). Its core role is to control the generation results of the diffusion model so that they meet specific conditions (such as style, content, motion, etc.). In one example, in video generation or time series data generation, the conditional frame can be a previous frame (such as images of the past few frames) to maintain temporal consistency. In image generation, it can be a reference image (such as a sketch, segmentation mask) to control the generated content.

[0113] In the disclosed embodiments, the video generation model includes multiple processing layers. Taking the DiT model as an example, the DiT model includes multiple layers with data processing capabilities (i.e., processing layers), each of which can be referred to as a block. Based on the principles of the DiT model, the DiT model can perform attention calculations to denoise image or video content (typically involving N steps), thereby generating a new video. During the calculation process, each block requires N steps to be executed. After each block is completed, video generation is achieved.

[0114] Among them, each block needs to query the token sequence of each conditional frame to perform attention calculation during the process of performing attention calculation on the noise feature sequence. Since the token sequence corresponding to the multimodal data of the time series type has a high degree of redundancy, based on this, when each block of the video generation module performs calculations, if it needs to query the token sequence of the conditional frame, it can filter the target conditional frame from the stored conditional frames according to the preset conditional frame screening strategy. Specifically, the token sequence of the target conditional frame can be selected from the token sequences corresponding to each conditional frame according to a preset ratio. Alternatively, the target conditional frame can be determined based on the number of layers of the current block, and then the token sequence of the target conditional frame can be queried. For example, if the number of layers of the current block is an odd number, the token sequence corresponding to the target conditional frame with an odd index is selected. Then, the block can perform attention calculation based on the feature sequence corresponding to the target conditional frame and the noise feature sequence.

[0115] In the above scheme, filtering the token sequence of the target condition frame for attention calculation can effectively reduce the number of tokens, optimize the efficiency of attention calculation, and thus improve the efficiency of video generation.

[0116] Optionally, each processing layer can also adopt a stride-number KV (Key-Value) cache reuse strategy to accelerate calculations. Specifically, when each block calculates the token sequence (i.e., context sequence) of time-series multimodal data, it can update the token sequence of the multimodal data only during the first step and store the updated token sequence. The subsequent N-1 steps do not update the token sequence, but instead call the stored token sequence for calculation, which can improve the efficiency of attention calculation. At the same time, when calculating the context sequence corresponding to each multimodal data, attention calculation can be performed only based on the context sequence corresponding to the multimodal data, without searching for the context sequences corresponding to other multimodal data. In this way, updates to the noise feature sequence will not affect the calculation results of the context sequence.

[0117] It can be understood that the above-mentioned dynamic token selection strategy and the stride number KV cache reuse strategy can be used together to improve the efficiency of attention calculation, thereby improving the efficiency of video generation.

[0118] In an exemplary embodiment, Figure 4 As shown in Figure 1, task information, visual description information, and content feature sequence are input into the video generation model to generate the target video, including:

[0119] In step S410, task information, visual description information, and content feature sequence are input into a video generation model.

[0120] In step S420, the task feature sequence corresponding to the task information is superimposed on the content feature sequence through the video generation model to obtain an updated content feature sequence.

[0121] In step S430, a target video is generated based on the updated content feature sequence, task information, and visual description information through a video generation model.

[0122] In an embodiment of the present disclosure, the video generation module may include an LTE (Learnable Task Embedding) module. The LTE module may learn a learnable task embedding vector (i.e., a task feature sequence) for each video generation task. The task feature sequence may be used to identify the video generation task. The LTE module and the video generation task may have a one-to-one correspondence. The video generation module may receive task information output by the intent recognition model and, based on the mapping relationship between the LTE module and the video generation task and the task identifier contained in the task information, determine the corresponding LTE module. The task feature sequence corresponding to the LTE module is then superimposed on the content feature sequence to obtain an updated content feature sequence. The target video is then generated by performing calculations based on the updated content feature sequence, the task information, and the visual description information.

[0123] In the above solution, the task feature sequence and content feature sequence are first superimposed before video generation. This strengthens the task characteristics and eliminates task ambiguity or unclear reference issues caused by multimodal input. For example, when a user uploads a video, they may have different purposes, such as wanting to replicate the camera movement (camera replication) or the character movements (action imitation). For the DiT model, if RoPE (Rotary Position Embedding) is shared, it will be difficult for the model to determine the user's intended task based on the uniform video conditions. Therefore, by introducing LTE to clarify the task, it can effectively eliminate ambiguity and improve the accuracy of video generation.

[0124] Figure 5This is a schematic diagram illustrating a framework of a video generation method according to an exemplary embodiment. The method specifically includes: an input module; a parsing module comprising an MLLM model and a PE model; and a content generation module comprising a noise video latent variable module, a contextual condition module, and a DiT model. During the execution of the video generation method, the input module can receive multimodal data input by the user and transmit the multimodal data to the parsing module and the contextual condition module, respectively. The MLLM model performs intent analysis based on the multimodal information to obtain task information for the video generation task, and transmits this task information to the PE model, the contextual condition module, and the DiT model, respectively. The PE model obtains visual description information for the video generation task based on the task information and multimodal data. The PE model then transmits this visual description information to the DiT model. The noise video latent variable module can generate a noise feature sequence and transmit it to the contextual condition module. The contextual condition module can generate a contextual condition control sequence based on the multimodal data and task information. The noise feature sequence and the contextual condition control sequence can be concatenated to obtain a content feature sequence, which is then transmitted to the DiT model. The DiT model then performs video generation calculations based on the task information, visual description information, and content feature sequences to produce the target video. The specific processing steps for each step can be found in the above descriptions and will not be further elaborated here. In this embodiment, the DiT model can adopt a contextual control paradigm. By tokenizing different control conditions and combining them with the noise feature sequence into a long token sequence, the entire token sequence can be jointly involved in the attention calculation. In a data-driven manner, the DiT model can flexibly utilize different contextual conditions to achieve various tasks, effectively improving its adaptability to different tasks and the generalization of video generation, allowing for efficient adaptation to diverse application scenarios. Furthermore, the above scheme utilizes a unified multimodal generation framework. Specifically, through a unified two-stage multimodal generation scheme (i.e., a multimodal data parsing stage and a video generation stage), users can flexibly combine multimodal signals such as commands, images, and videos to achieve various multimodal control tasks. This unified framework enhances task compatibility and reduces parameter redundancy. Furthermore, the specific framework adopts a modular design, resulting in a simple structure and strong interchangeability.

[0125] It should be understood that although Figure 1-Figure 4 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1-Figure 4At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0126] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can be referred to each other, and each embodiment focuses on the differences from other embodiments. For related parts, please refer to the description of other method embodiments.

[0127] Figure 6 FIG. 1 is a block diagram of a video generation device according to an exemplary embodiment. Figure 6 , the device includes an acquisition unit 610, a first determination unit 620, a first generation unit 630 and a second generation unit 640.

[0128] According to a second aspect of an embodiment of the present disclosure, there is provided a video generating apparatus, including:

[0129] An acquiring unit 610 is configured to execute acquiring a video generation request, wherein the video generation request includes multimodal data input by a user;

[0130] A first determining unit 620 is configured to determine task information of a video generation task corresponding to the video generation request and visual description information of the video generation task based on the multimodal data;

[0131] A first generating unit 630 is configured to generate a context condition control sequence based on the multimodal data, and generate a content feature sequence based on the context condition control sequence;

[0132] The second generating unit 640 is configured to input the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video.

[0133] In one embodiment, the first determining unit 620 includes:

[0134] A first determining subunit is configured to input the multimodal data into an intent recognition model to obtain task information of a video generation task corresponding to the video generation request;

[0135] The second determining subunit is configured to input the task information and the multimodal data into a content expansion model to obtain visual description information of the video generation task.

[0136] In one embodiment, the multimodal data includes base content data and data indicating a spatiotemporal position; the first generating unit 630 includes:

[0137] an encoding subunit configured to perform encoding processing on the base content data and the data indicating the spatiotemporal position, respectively, to obtain a first encoding sequence corresponding to the base content data and a second encoding sequence corresponding to the data indicating the spatiotemporal position;

[0138] The first superposition subunit is configured to superpose the first coding sequence and the second coding sequence to obtain a context condition control sequence.

[0139] In one embodiment, the first superposition subunit is configured to perform:

[0140] The codes at the same position in the first coding sequence and the second coding sequence are superimposed to obtain a context condition control sequence.

[0141] In one embodiment, the first generating unit 630 includes:

[0142] The splicing subunit is configured to splice the context condition control sequence and the noise feature sequence to obtain a content feature sequence.

[0143] In one embodiment, the multimodal data includes time series multimodal data and / or non-time series multimodal data.

[0144] In one embodiment, the apparatus further comprises:

[0145] A second determining unit is configured to determine a target index range for a target sequence in the content feature sequence based on attribute information corresponding to the target sequence; wherein the target sequence is a feature sequence corresponding to non-time series multimodal data;

[0146] An allocating unit is configured to allocate index information to the target sequence based on the target index range.

[0147] In one embodiment, the apparatus further comprises:

[0148] a screening unit configured to execute, for each processing layer of the video generation model, a process of performing attention calculation on the noise feature sequence through the processing layer, and to screen a target conditional frame from the stored conditional frames according to a preset conditional frame screening strategy;

[0149] A computing unit is configured to perform attention calculation based on the feature sequence corresponding to the target condition frame and the noise feature sequence.

[0150] In one embodiment, the second generating unit 640 includes:

[0151] an input subunit, configured to input the task information, the visual description information, and the content feature sequence into a video generation model;

[0152] A second superposition subunit is configured to execute, through the video generation model, superimposing the task feature sequence corresponding to the task information and the content feature sequence to obtain an updated content feature sequence;

[0153] The generation subunit is configured to generate a target video based on the updated content feature sequence, the task information and the visual description information through the video generation model.

[0154] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0155] Figure 7 FIG2 is a block diagram of an electronic device 700 for a video generation method according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0156] Reference Figure 7 The electronic device 700 may include one or more of the following components: a processing component 702 , a memory 704 , a power supply component 706 , a multimedia component 708 , an audio component 710 , an input / output (I / O) interface 712 , a sensor component 714 , and a communication component 716 .

[0157] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 702 may include one or more modules to facilitate interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate interaction between the multimedia component 708 and the processing component 702.

[0158] The memory 704 is configured to store various types of data to support operations on the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, or graphene memory.

[0159] The power supply component 706 provides power to the various components of the electronic device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 700.

[0160] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the electronic device 700 is in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.

[0161] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 also includes a speaker for outputting audio signals.

[0162] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0163] The sensor assembly 714 includes one or more sensors for providing various aspects of status assessment for the electronic device 700. For example, the sensor assembly 714 can detect the open / closed state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor assembly 714 can also detect changes in the position of the electronic device 700 or components of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the device 700, and temperature changes of the electronic device 700. The sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 714 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0164] The communication component 716 is configured to facilitate wired or wireless communication between the electronic device 700 and other devices. The electronic device 700 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0165] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0166] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions. The instructions can be executed by the processor 720 of the electronic device 700 to perform the above method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0167] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions, and the instructions can be executed by the processor 720 of the electronic device 700 to implement the above method.

[0168] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. can also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments and will not be described one by one here.

[0169] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0170] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video generation method, characterized in that: include: Obtaining a video generation request, where the video generation request includes multimodal data input by a user; Determining, based on the multimodal data, task information of a video generation task corresponding to the video generation request and visual description information of the video generation task; generating a context condition control sequence based on the multimodal data, and generating a content feature sequence based on the context condition control sequence; Inputting the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video; The multimodal data includes base content data and spatiotemporal position indication data, wherein the spatiotemporal position indication data is used to indicate the spatial position and / or temporal position of a target object, the base content data is basic content for generating the target video, and the spatiotemporal position indication data and the base content data are multimodal data of a time-series type; The generating of the context condition control sequence based on the multimodal data includes: encoding the base content data and the indication data of the spatiotemporal position respectively to obtain a first encoding sequence corresponding to the base content data and a second encoding sequence corresponding to the indication data of the spatiotemporal position; and superimposing the first encoding sequence and the second encoding sequence to obtain a context condition control sequence.

2. The video generation method according to claim 1, wherein: The determining, based on the multimodal data, task information of the video generation task corresponding to the video generation request and visual description information of the video generation task includes: Inputting the multimodal data into an intent recognition model to obtain task information of a video generation task corresponding to the video generation request; The task information and the multimodal data are input into a content expansion model to obtain visual description information of the video generation task.

3. The video generation method according to claim 1, wherein: The step of superimposing the first coding sequence and the second coding sequence to obtain a context condition control sequence includes: The codes at the same position in the first coding sequence and the second coding sequence are superimposed to obtain a context condition control sequence.

4. The video generation method according to claim 1, wherein: The generating of the content feature sequence based on the context condition control sequence includes: The context condition control sequence and the noise feature sequence are concatenated to obtain a content feature sequence.

5. The video generation method according to any one of claims 1 to 4, characterized in that: The multimodal data includes time series multimodal data and / or non-time series multimodal data.

6. The video generation method according to claim 5, characterized in that: After inputting the task information, the visual description information, and the content feature sequence into the video generation model, the method further includes: For a target sequence in the content feature sequence, determining a target index range based on attribute information corresponding to the target sequence; wherein the target sequence is a feature sequence corresponding to non-time series multimodal data; Based on the target index range, index information is assigned to the target sequence.

7. The video generation method according to claim 1, characterized in that: The video generation model includes multiple processing layers; after inputting the task information, the visual description information, and the content feature sequence into the video generation model, the method further includes: For each processing layer of the video generation model, in the process of performing attention calculation on the noise feature sequence through the processing layer, a target conditional frame is selected from the stored conditional frames according to a preset conditional frame selection strategy; Attention calculation is performed based on the feature sequence corresponding to the target condition frame and the noise feature sequence.

8. The video generation method according to claim 1, wherein: The step of inputting the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video includes: Inputting the task information, the visual description information, and the content feature sequence into a video generation model; superimposing the task feature sequence corresponding to the task information and the content feature sequence through the video generation model to obtain an updated content feature sequence; A target video is generated by the video generation model based on the updated content feature sequence, the task information and the visual description information.

9. A video generating device, characterized in that: include: an acquiring unit configured to execute acquiring a video generation request, wherein the video generation request includes multimodal data input by a user; A first determining unit is configured to determine task information of a video generation task corresponding to the video generation request and visual description information of the video generation task based on the multimodal data; A first generating unit is configured to generate a context condition control sequence based on the multimodal data, and generate a content feature sequence based on the context condition control sequence; A second generating unit is configured to input the task information, the visual description information, and the content feature sequence into a video generation model to generate a target video; The multimodal data includes base content data and spatiotemporal position indication data, wherein the spatiotemporal position indication data is used to indicate the spatial position and / or temporal position of a target object, the base content data is basic content for generating the target video, and the spatiotemporal position indication data and the base content data are multimodal data of a time-series type; The first generation unit includes: an encoding sub-unit, configured to perform encoding processing on the base content data and the indication data of the spatiotemporal position respectively, to obtain a first encoding sequence corresponding to the base content data, and a second encoding sequence corresponding to the indication data of the spatiotemporal position; a first superposition sub-unit, configured to perform superposition of the first encoding sequence and the second encoding sequence to obtain a context condition control sequence.

10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the video generation method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generating method according to any one of claims 1 to 8.