Video generation method and system based on large-scale cultural video model
By introducing spatial position and grammatical perceptual contrast constraints in the large-scale literary video model, the latent spatial noise graph sequence is optimized, and the problem of motion misalignment of multiple subjects is solved, and the alignment of video generation and text description is achieved.
Patent Information
- Application Number
- CN202410948632.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-07-16
AI Technical Summary
When the existing large-scale Wensheng video model generates multiple objects with different actions, it is often difficult to accurately understand the number of subjects and the binding of action, resulting in the generated videos that do not match the text description, especially in the ambiguous motion of multiple subjects.
By introducing spatial position constraints and grammatical perceptual contrast constraints, using the dual attention of space and grammatical for guidance, the latent space noise diagram sequence is optimized to ensure the motion alignment of multiple subjects.
The multi-subject and motion correspondence of the existing T2V model is enhanced, the generated video has a higher consistency with the text description, and the subject and action are correctly bound.
Smart Images

Figure CN119031209B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of diffusion model T2V generation in the field of artificial intelligence technology, and more specifically, to a method and system for generating videos with multiple subjects participating in different actions based on a large-scale cultural video model and guided by spatial and grammatical dual attention. Background Art
[0002] Text-to-Video (T2V) technology is an innovative advancement in artificial intelligence. By combining natural language processing, image processing, and machine learning, it achieves breakthroughs in converting text descriptions into video content. It holds broad application prospects in fields such as education, healthcare, security, and entertainment. In recent years, text-to-image generation models based on diffusion models have made significant progress, capable of generating high-quality images that closely correlate with text descriptions. Building on this progress, researchers have expanded the application of diffusion models to text-to-video generation.
[0003] However, existing large-scale text-based video models primarily focus on simple scenes with a single object and action. Therefore, when users need to generate multiple objects with different actions, these models often struggle to accurately understand the number of subjects in the prompt, leading to a mismatch in the number of subjects and misaligned prompts, meaning that the generated video does not match the semantics of the text description. A major issue is the misalignment between objects and their corresponding actions. For example, the prompt "a person skateboarding, a dog sitting" may result in an incorrect number of objects, with two dogs or only one person, or incorrect action binding, with both the person and the dog skateboarding, or the person sitting and the dog skateboarding.
[0004] Therefore, how to solve the problem of uneven motion of multiple subjects in video generation tasks has become one of the current research directions in the field of artificial intelligence video generation. Summary of the Invention
[0005] In view of the problem of misaligned motion of multiple subjects in the above-mentioned current Vincent video solutions, the purpose of the present invention is to provide a video generation method and system based on a large-scale Vincent video model. For multiple subjects, especially multiple subjects with different actions, by introducing spatial position constraints and grammatical perception contrast constraints, the dual attention of space and grammar is guided to align the motion of the generated video, thereby enhancing the multi-subject and motion correspondence of the existing T2V model.
[0006] In one aspect, the present invention provides a video generation method based on a large-scale cultural video model, which generates videos for multiple subjects, comprising the following steps:
[0007] S110: randomly generating a latent space noise map sequence based on a preset large-scale Wensheng video model;
[0008] S120: Denoising the latent space noise map sequence based on the Vincent video input data, and optimizing the latent space noise map sequence within a preset time step in the denoising process using a spatial position constraint and a grammatical-aware contrast constraint; wherein the Vincent video input data includes a text description of the video to be generated, the text description including multiple subjects in the video to be generated and their corresponding actions and background, and bounding boxes of spaces corresponding to the multiple subjects in the video to be generated;
[0009] S130: decoding the latent space noise map sequence obtained in the last step of the denoising process to generate a video;
[0010] The optimization processing of the latent space noise map sequence includes:
[0011] S121: Inputting the latent space noise map sequence of the current step and the text description of the video to be generated into a preset noise prediction network to perform noise prediction;
[0012] S122: Extracting a cross-attention graph sequence generated in the noise prediction network;
[0013] S123: Calculating spatial position constraints and grammatical perception contrast constraints using the cross attention map sequence;
[0014] S124: Calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint with respect to the latent space noise map sequence through back propagation, and update the latent space noise map sequence using the gradients.
[0015] Among them, an optional solution is to use the spatial position constraints calculated by the cross-attention map sequence, including spatial perception constraints and background constraints;
[0016] The spatial perception constraint is calculated based on the bounding boxes of the spaces corresponding to the multiple subjects in the generated video in the Wensheng video input data. The loss function of the spatial perception constraint is as follows:
[0017]
[0018] Among them, Λ fg is the loss function of spatial perception constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated;
[0019] The loss function of the background constraint is as follows:
[0020]
[0021] Among them, Λ bg is the loss function of the background constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated;
[0022] The loss function of the spatial position constraint is as follows:
[0023] Λ sp =λ fg Λ fg +λ bg Λ bg
[0024] Among them, Λ sp is the loss function of spatial position constraint, λ fg is the loss function of spatial perception constraint Λ fg The weight coefficient, λ bg is the loss function Λ for background constraints bg The weight coefficient of .
[0025] Among them, an optional solution is that the use of the cross-attention map sequence to calculate the grammatical perception contrast constraint includes:
[0026] The text description S in the video input data * The word pair consisting of the i-th word representing the subject name and the j-th word describing the subject action As positive samples of each other, the text representation S in the video input data * Other word sets in U i As negative samples;
[0027] The forward loss function of the grammatical-aware contrast constraint is used to minimize the s in the cross-attention map i and s j The distance between the mappings is expressed as:
[0028]
[0029] in, is the forward loss function of the grammatically aware contrastive constraint, f dist (·) is the distance function calculation between the cross attention maps, and They are the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word and the j-th word in , where F is the number of image frames in the video to be generated;
[0030] The negative loss function of the grammatical-aware contrast constraint is used to encourage the cross-attention map of the subject and its action word pair to stay away from S * The cross attention map of words in the other word set U in is expressed as:
[0031]
[0032] in, represents the negative loss function of the grammatical-aware contrast constraint, f dist (·) represents the distance function calculation between cross-attention maps, and They are the text description S in the Wensheng video input data * The f-th frame cross attention map corresponding to the i-th word and the u-th word in , where F is the number of image frames in the video to be generated;
[0033] The loss function of the grammar-aware contrast constraint is calculated as follows:
[0034]
[0035] Among them, Λ syt Loss function representing the syntax-aware contrastive constraint.
[0036] Among them, an optional scheme is that the use of the gradient to update the latent space noise map sequence includes: in the early stage of denoising, subtracting the gradient of the cross-attention constraint based on the spatial position from each latent space noise map in the latent space noise map sequence at a preset learning rate; in the latent stage of denoising, subtracting the gradient of the grammatical perceptual contrast constraint from each latent space noise map in the latent space noise map sequence at a preset learning rate.
[0037] An optional solution is to calculate the gradient to update z at each time step t , as shown below:
[0038]
[0039] Among them, α represents the learning rate of the optimization process, λ * Represents the weight of the constraint.
[0040] On the other hand, the present invention further provides a video generation system based on a large-scale Wensheng video model, which generates videos using the aforementioned video generation method based on a large-scale Wensheng video model, including:
[0041] A latent space noise map sequence acquisition unit, configured to randomly generate a latent space noise map sequence based on a preset large-scale Wensheng video model;
[0042] a denoising unit configured to perform denoising processing on the latent space noise map sequence based on Vincent video input data, wherein the Vincent video input data includes a text description of the video to be generated, the text description including a plurality of subjects in the video to be generated and their corresponding actions and background, and bounding boxes of spaces corresponding to the plurality of subjects in the video to be generated;
[0043] an optimization unit, configured to optimize the latent space noise map sequence within a preset time step in the denoising process using spatial position constraints and grammatical perception contrast constraints;
[0044] A decoding unit, configured to decode the latent space noise map sequence obtained in the last step of the denoising process to generate a video;
[0045] Wherein, the optimization unit includes:
[0046] A noise prediction unit is used to input the latent space noise map sequence of the current step and the text description of the video to be generated into a preset noise prediction network to perform noise prediction;
[0047] a cross-attention extraction unit, configured to extract a cross-attention graph sequence generated in the noise prediction network;
[0048] an attention adjustment constraint calculation unit, configured to calculate spatial position constraints and grammatical perception contrast constraints using the cross-attention map sequence;
[0049] An updating unit is configured to calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint on the latent space noise map sequence through back propagation, and update the latent space noise map sequence using the gradients.
[0050] The present invention further provides an electronic device, comprising:
[0051] at least one processor; and,
[0052] a memory communicatively connected to the at least one processor; wherein,
[0053] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the steps in the video generation method based on the large-scale cultural video model as described above.
[0054] From the above technical solution, it can be seen that the video generation method and system based on the large-scale text-generated video model provided by the present invention, by introducing attention-based constraints based on spatial position and grammatical perception, uses spatial and grammatical dual attention to guide the large-scale diffusion text-generated video model to solve the problem of motion disorder of multiple subjects, especially multiple subjects with different actions; wherein, the latent space noise map is first optimized within a preset time step of the latent space noise map denoising process, and the optimization process includes: first, applying cross-attention constraints based on spatial position in the early stage of the denoising process so that multiple nouns can clearly focus on the correct object area; then using the cross-attention map sequence to calculate the grammatical perception-based contrast constraint to improve the intersection of verbs and corresponding nouns, and promote the correct binding of action subjects; finally, the gradient of the spatial position constraint and the grammatical perception contrast constraint to the latent space noise map sequence is calculated through back propagation, and the latent space noise map sequence is updated using this gradient, so that the generated video achieves motion alignment and enhances the multi-subject and motion correspondence of the existing T2V model. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] By referring to the following description in conjunction with the accompanying drawings, and with a more complete understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the accompanying drawings:
[0056] Figure 1 Flowchart of a video generation method based on a large-scale cultural video model according to an embodiment of the present invention;
[0057] Figure 2 2 is a schematic diagram showing the principle of a video generation method based on a large-scale cultural video model according to an embodiment of the present invention;
[0058] Figure 3 Schematic diagram of a framework of a system based on a large-scale cultural video model according to an embodiment of the present invention;
[0059] Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present invention;
[0060] Figure 5 and Figure 6 They are respectively example diagrams of the effects of applying the video generation method and system based on the large-scale cultural video model of the present invention to generate videos. DETAILED DESCRIPTION
[0061] In the following description, for illustrative purposes, numerous specific details are set forth to provide a comprehensive understanding of one or more embodiments. However, it will be apparent that the embodiments may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to facilitate description of one or more embodiments.
[0062] In response to the problems of multiple subjects and motion misalignment in existing Vincent video solutions, the present invention provides a video generation method and system based on a large-scale Vincent video model. For multiple subjects, especially multiple subjects with different actions, based on the descriptive text of the multiple subjects in the to-be-generated video and their corresponding actions and backgrounds, as well as the bounding boxes of the multiple subjects in the generated video in space, the latent space noise map sequence and text description of the current step are input into a noise prediction network; then the cross-attention map sequence is extracted, and the cross-attention map sequence is used to calculate the spatial position constraint and the contrast constraint based on grammar perception; finally, the gradient of the attention adjustment constraint to the latent space noise map sequence is calculated through backpropagation, and the latent space noise map sequence is updated using this gradient, so that the generated video achieves the effect of alignment of the subject and action, while ensuring the consistency of the generated video with the descriptive text.
[0063] In order to better illustrate the technical solution of the present invention, some technical terms involved in the present invention are briefly explained below.
[0064] A text-driven video model generates a video from a simple text description (in English). Existing text-driven video models include AnimateDiff and ModelScope. They all use the diffusion model framework for text-driven generation.
[0065] The noise prediction network uses a transformer-based backbone network to process inputs of different modalities to predict noise. For example, a U-net network is used to predict noise at each step. U-net is a semantic segmentation algorithm based on a fully convolutional network that utilizes a symmetrical U-shaped structure and mirroring operations to process boundary pixels.
[0066] Cross attention is a mechanism used in the noise prediction U-net network structure of ModelScope. The idea is to enable one sequence to "pay attention" to another sequence.
[0067] The back-propagation (BP) algorithm calculates the gradient of the loss function with respect to each parameter in the neural network through the derivative chain rule, and updates the parameters in conjunction with the optimization method to reduce the loss function.
[0068] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0069] It should be noted that the following description of exemplary embodiments is merely illustrative and is not intended to limit the present invention, its application, or use. Techniques and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques and devices should be considered part of the specification.
[0070] To illustrate the video generation method and system based on the large-scale cultural video model provided by the present invention, Figure 1 、 Figure 2 The process and principle of the video generation method based on the large-scale cultural video model according to an embodiment of the present invention are exemplarily marked respectively; Figure 3 The logical structure of the video generation system based on the large-scale cultural video model according to the embodiment of the present invention is exemplarily indicated.
[0071] like Figure 1 and Figure 2 As shown in the figure, the video generation method based on the large-scale cultural video model provided by the present invention generates videos for multiple subjects, especially multiple subjects with different actions, and mainly includes the following steps:
[0072] S110: randomly generating a latent space noise map sequence based on a preset large-scale Wensheng video model;
[0073] S120: Denoising the latent space noise map sequence based on the Vincent video input data, and optimizing the latent space noise map sequence within a preset time step in the denoising process using a spatial position constraint and a grammatical-aware contrast constraint; wherein the Vincent video input data includes a text description of the video to be generated, the text description including multiple subjects in the video to be generated and their corresponding actions and background, and bounding boxes of spaces corresponding to the multiple subjects in the video to be generated;
[0074] S130: Decode the latent space noise map sequence obtained in the last step of the denoising process to generate a video.
[0075] Specifically, as an example, the large-scale Vincent video model based on the present invention can adopt the diffusion model ModelScope. For the convenience of description, in the following embodiments, ModelScope is used as an implementation example of the large-scale Vincent video model.
[0076] The following will provide a more detailed description of the above steps of the video generation method based on a large-scale cultural video model with dual spatial and grammatical attention guidance in conjunction with specific embodiments.
[0077] For step S110, the latent space noise map sequence randomly generated based on the preset large-scale Wensheng video model is represented as z T , the latent space noise map sequence is a randomly generated input that obeys Gaussian distribution.
[0078] In step S120, the latent space noise map sequence is denoised based on the Vincent video input data, and the latent space noise map obtained by the denoising process is decoded to generate a video; the first few steps of the denoising process in S120 need to be optimized first, and then the denoising process is performed.
[0079] In order to realize the generation of multiple subject videos with different actions, the present invention optimizes the latent space noise map sequence by utilizing spatial position constraints and grammatical perception contrast constraints within a preset time step in the denoising process.
[0080] More specifically, in one embodiment of the present invention:
[0081] Assuming that the user inputs the video data as: "A man is skateboarding and a dog is sitting", then the latent space noise map sequence is represented as z T Afterwards, the pre-trained text encoder encodes the received user input text "A man is skateboarding and a dog is sitting" into an embedding vector and injects it into the U-Net network of ModelScope;
[0082] For the randomly generated latent space noise map sequence z T The embedded vector of the text description will serve as conditional information to guide the U-Net network to perform the denoising process. The U-Net network calculates the corresponding predicted noise sequence through the latent space noise map sequence obtained in the previous step and the embedded vector of the text description. Then, the noise map sequence of this step is obtained by subtracting the predicted noise sequence from the latent space noise map sequence obtained in the previous step.
[0083] Furthermore, the latent space noise map sequence is optimized within a preset time step of the denoising process. The U-Net is a neural network consisting of a decoder, an encoder, and a bottleneck layer. The encoder extracts features by reducing the feature map through convolution and pooling. The decoder is symmetrical to the encoder, performing upsampling followed by concatenation of feature maps with the encoder to preserve features before performing convolution. The bottleneck layer consists of two 3×3 convolutional layers and one 1×1 convolutional layer; the 1×1 convolutional layer is used for the final output.
[0084] In one embodiment of the present invention, the denoising process is performed in 50 time steps, and the optimization process for the latent space noise map sequence is performed in 25 time steps. That is, the denoising process is repeated for 50 steps, and an optimization process for the latent space noise map sequence is added within the [50, 25] time step of the denoising process. The optimization process is performed first, and then the denoising process is performed after the optimization process is completed. This effectively removes noise from the latent space noise map sequence and generates a video that better matches the text description entered by the user.
[0085] The optimization process proposed in the present invention is equivalent to the denoising process at t = [50, 0] being a large framework of the diffusion model. Within this large framework, the latent space noise map sequence is optimized before the denoising process at t = [50, 25]. Specifically, the optimization process of the latent space noise map sequence includes:
[0086] S121: embedding the text obtained by encoding the latent space noise map sequence of the current step and the text description of the video to be generated into a preset noise prediction network to perform noise prediction; the raw video input data includes the text description of the video to be generated, the text description including multiple subjects in the video to be generated and their corresponding actions and background, and bounding boxes of the space corresponding to the multiple subjects in the video to be generated, for generating a video based on the text description;
[0087] S122: Extracting a cross-attention graph sequence generated in the noise prediction network;
[0088] S123: Calculating spatial position constraints and grammatical perception contrast constraints using the cross attention map sequence;
[0089] S124: Calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint with respect to the latent space noise map through back propagation, and use the gradients to update the latent space noise map; wherein, in the process of using the gradients to update the latent space noise map, the spatial position constraint denoising update is used in the early stage, and the grammatical perception contrast constraint denoising update is used in the later stage.
[0090] In a specific embodiment of the present invention, the above optimization process is executed 25 times in a loop. In the first 5 times, spatial position constraints are executed, and in the 5th to 25th times, grammar-aware contrast constraints are executed.
[0091] More specifically, as an example, the optimization process for the latent space noise map sequence includes the following operations:
[0092] The latent space noise map sequence z of the current step t The text description P of the video to be generated is embedded into the U-Net network;
[0093] Extract the five 16×16 cross-attention map sequences generated by the U-Net network;
[0094] Computing spatial position constraints and grammatically aware contrast constraints using the cross-attention map;
[0095] The gradients of the spatial position constraint and the grammatical perception contrast constraint on the latent space noise map sequence are calculated by back propagation, and the latent space noise map sequence is updated using the gradients. The spatial position constraint denoising is used to update the denoising in the early stage (for example, the first 5 steps of denoising), and the grammatical perception contrast constraint denoising is used to update the denoising in the late stage (for example, the 5th to 25th steps of denoising).
[0096] After denoising, the decoder is used to decode the latent space sequence noise map sequence z0 obtained in the last step into a 16-frame video with a resolution of 512×512.
[0097] exist Figure 1 、 Figure 2 In the embodiments shown together, spatial position constraints and grammatical perception contrast constraints (collectively referred to as attention adjustment constraints) are integrated into a large-scale cultural video model, and spatial position constraints and grammatical perception contrast constraints are introduced in the denoising process without the need for additional training.
[0098] When introducing spatial position constraints and grammatical perception contrast constraints in the denoising process, at each time step t, the potential noise is first input into the noise prediction network, and all 16×16 resolution cross-attention maps are extracted. Then, the spatial position constraints and grammatical perception contrast constraints are calculated using the cross-attention maps corresponding to each subject and subject action.
[0099] Specifically, as an example, the text description P input by the user is defined as a word set S = {s1, s2, ..., s L}, where L is the length of the text description P. This embodiment defines a set of word pairs consisting of nouns and verbs representing multiple subjects and corresponding actions in P, which is expressed as s i It is a noun that indicates the name of the subject. j is the verb describing the subject's action, i and j are the index positions of the subject noun and action verb in P. In order to solve the problem of the lack of focus of the cross-attention map of nouns in the early stage of denoising, this paper introduces an additional spatial prior, namely the bounding box, to assist the denoising process, with the aim of guiding the subject's nouns to focus on the specified area and separate from each other. S * The spatial location prior set corresponding to the subject noun in is represented as in, represents the spatial position prior of the fth frame corresponding to the i-th word in S, Represents the cross attention map of the fth frame corresponding to the i-th word in S. In order to solve the problem that there is no clear connection between the subject and the corresponding action, which leads to the misalignment of the subject and the corresponding action, this paper introduces the grammatical perception contrast constraint and establishes S * There is a clear connection between the subject noun and the corresponding action verb.
[0100] In other words, the spatial position constraint and grammatical perceptual contrast constraint are calculated using the cross-attention map obtained from the subject nouns and the corresponding action verbs. Subject nouns refer to the nouns representing the subject in the text prompt entered by the user, and action verbs refer to the verbs representing the subject's actions in the text prompt entered by the user. For example, in the user's input text "A man is skateboarding and a dog is sitting," the subject nouns are "man" and "dog," and the corresponding action verbs are "skateboarding," representing the action of "man," and "sitting," representing the action of "dog."
[0101] In a specific embodiment of the present invention, the spatial position constraint ensures that the position and quantity of the subject represented by the subject noun conform to the description of the input text, and the grammatical perception contrast constraint establishes the connection between the subject noun and the corresponding action verb, ensuring that the subject's corresponding action can only appear on the subject.
[0102] Specifically, in order to ensure that the cross-attention map of nouns is focused on the area defined by a given spatial prior, this paper proposes a spatial position constraint, which consists of a spatial perception constraint and a background constraint, to enhance the focus of the cross-attention map sequence on foreground objects. The loss function of the spatial perception constraint is as follows:
[0103]
[0104] Among them, Λ fg is the loss function of spatial perception constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, F is the number of image frames of the video to be generated.
[0105] In order to prevent nouns from considering information outside the corresponding bounding box, thereby generating multiple entities represented by nouns outside the specified area, this paper proposes background constraints to reduce the influence of the cross-attention map of nouns outside the bounding box. The loss function of the background constraint is as follows:
[0106]
[0107] Among them, Λ bg is the loss function for background constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, F is the number of image frames of the video to be generated.
[0108] Accordingly, the complete spatial position constraint is defined as:
[0109] Λ sp =λ fg Λ fg +λ bg Λ bg
[0110] Among them, Λ sp is the loss function of spatial position constraint, λ fg is the loss function of spatial perception constraint Λ fg The weight coefficient, λ bg is the loss function Λ for background constraints bg The weight coefficient of .
[0111] Furthermore, considering the verb s j Cross attention map Should be aligned j Corresponding noun s i For example, the cross attention map of the verb "skateboarding" should be in the area corresponding to the noun "man". Therefore, the present invention also applies the above spatial position constraint to s j The cross-attention map uses the bounding boxes of the corresponding nouns to enforce this alignment.
[0112] Under the spatial position constraints of nouns and verbs, the cross-attention maps of verbs can be aligned to the regions defined by the spatial prior to a certain extent. However, there is still a lack of clear relationship between the cross-attention maps of verbs and nouns, resulting in the leakage of actions to other subjects, for example, the action "skateboarding" of the subject "man" appears on the subject "dog". Therefore, the present invention proposes to use syntactic relations to establish a strong connection between verbs and nouns, thereby enhancing the alignment between the movement and the corresponding subject. To achieve this goal, the present invention introduces contrastive learning to minimize the distance between the cross-attention map of the verb and the cross-attention map of the corresponding noun, while keeping it away from the cross-attention maps of other words to prevent interference.
[0113] Based on this, the present invention proposes a grammatical-aware contrast constraint, in which each The subject noun and the corresponding action verb pair in S are used as each other’s positive samples, while the other words in S are used as their negative samples. That is, the text description S in the video input data is used as the negative sample of each other. * The i-th word representing the subject name (noun S i ) and the jth word describing the subject action (noun S j ) As positive samples of each other, the text representation S in the video input data * Other word sets in U i as negative samples.
[0114] right The subject noun s in i and action verbs j , the grammatically aware contrast-constrained forward loss aims to minimize the cross-attention map s i and s j The distance between the cross attention maps. The syntax-aware contrast constraint loss function is as follows:
[0115]
[0116] in, represents the forward loss function of the grammatical-aware contrast constraint, f dist (·) represents the calculation of the distance function between cross-attention maps, and They are the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word and the j-th word in , where F is the number of image frames in the video to be generated.
[0117] In addition, the present invention also proposes a grammatical perception contrast constraint negative loss function to encourage the cross attention graph of the subject and its action word pair to stay away from S * The cross attention map of words in the set U of other words in . Specifically, U is the set of other words for each subject; for The present invention defines the set of other words in S as U i , that is, U i is a set of other words representing the i-th word of the current subject. i Belongs to U. For example, in "a man is skateboarding and adog is sitting", the second word man (a noun indicating the name of the subject) and the U corresponding to the verb skateboarding i=[a,is,and,a,dog,is,sitting]; and the seventh word dog (a noun indicating the name of the subject) and the corresponding verb sitting U i =[a,man,is,skateboarding,and,a,is]. And U=[U i ,U i ]=[[a,is,and,a,dog,is,sitting],[a,man,is,skateboarding,and,a,is]].
[0118] The negative loss function of the grammatical-aware contrast constraint is used to keep the cross-attention map of the subject noun and its corresponding verb away from the attention map of the words in U. That is, if P is "a man is skateboarding and a dog issitting", then the cross-attention map of the subject noun "man" and the corresponding action verb "skateboarding" should be separated from the cross-attention maps of other words in P. The negative loss function of the grammatical-aware contrast constraint is as follows:
[0119]
[0120] in, represents the negative loss function of the grammatical-aware contrast constraint, f dist (·) represents the distance function calculation between cross-attention maps, and They are the text description S in the Wensheng video input data * The f-th frame cross attention map corresponding to the i-th word and the u-th word in , where F is the number of image frames in the video to be generated.
[0121] In summary, the complete grammatically aware contrast constraint loss function is expressed as follows:
[0122]
[0123] Among them, Λ syt Loss function representing the syntax-aware contrastive constraint.
[0124] That is to say, for the input text "Aman is skateboarding and a dog is sitting", the spatial position constraint ensures the position and number of the subjects "man" and "dog", and the grammatical perception contrast constraint aligns the action "skateboarding" with the subject "man" and the action "sitting" with "dog". In the early stage of denoising (such as the first 5 steps), the present invention calculates the gradient of the spatial position constraint on the latent space noise sequence in this step by back propagation, and then uses the obtained gradient to update the latent space noise map sequence. In the later stage of denoising (such as steps 5 to 25), the gradient of the grammatical perception contrast constraint on the latent space noise sequence in this step is calculated by back propagation, and then uses the obtained gradient to update the latent space noise map sequence. Among them, the formula for updating the latent space noise map sequence is as follows:
[0125]
[0126] In the above formula, the gradient of each noise map in the latent space noise map sequence is subtracted at a preset learning rate to obtain the updated latent space noise map sequence. * is the constraint weight, i.e. the scaling factor for gradient update, To calculate the gradient of the latent space noise map sequence obtained by back propagation, t is the current time step and t0 is the number of control update steps.
[0127] Preferably, at each time step t, z t It is updated iteratively until the preset number of update steps is reached.
[0128] In the early stages of the optimization process (e.g., the first 5 steps), the gradient of the spatial position constraint is calculated to guide the latent space noise map sequence to align with the specified spatial prior, accurately generate the correct number of subjects, and ensure that their movements correspond to the positions of the relevant objects. In the later stages (e.g., steps 5 to 25), the grammatical perception contrast constraint is calculated to further improve the latent space noise map sequence z t ,enhancing high-response attention alignment between nouns and verbs.
[0129] like Figure 3 As shown, the present invention also provides a video generation system based on a large-scale Vincent video model with dual spatial and grammatical attention guidance, which uses the video generation method based on a large-scale Vincent video model as described above to generate videos for multiple subjects, especially multiple subjects with different actions. The video generation 300 based on a large-scale Vincent video model with dual spatial and grammatical attention guidance mainly includes three parts: a latent space noise map acquisition unit 310, an optimization unit 320, a denoising unit 330 and a decoding unit 340.
[0130] The latent space noise map sequence acquisition unit 310 is used to randomly generate a latent space noise map sequence based on a preset large-scale Wensheng video model;
[0131] a denoising unit 320 configured to perform denoising on the latent space noise map sequence based on Vincent video input data, wherein the Vincent video input data includes a text description of the video to be generated, the text description including a plurality of subjects in the video to be generated and their corresponding actions, a background, and bounding boxes of spaces corresponding to the plurality of subjects in the video to be generated;
[0132] An optimization unit 330 is configured to optimize the latent space noise map sequence by using spatial position constraints and grammatical perception contrast constraints within a preset time step in the denoising process performed by the denoising unit;
[0133] The decoding unit 340 is configured to decode the latent space noise map sequence obtained in the last step of the denoising process to generate a video.
[0134] The optimization unit 330 includes:
[0135] The noise prediction unit 331 is used to embed the latent space noise map of the current time step and the text description of the video to be generated into a preset noise prediction network for noise prediction;
[0136] a cross-attention extraction unit 332 for extracting a cross-attention graph sequence generated in the noise prediction network;
[0137] an attention adjustment constraint calculation unit 333 for calculating spatial position constraints and grammatical perception contrast constraints using the cross attention map sequence;
[0138] An updating unit 334 is configured to calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint on the latent space noise map through back propagation; wherein, in the early stage of denoising, the spatial position constraint gradient is used to update the latent space noise map, and in the latent space noise map in the latent stage of denoising, the grammatical perception contrast constraint gradient is used to update the latent space noise map.
[0139] The attention adjustment constraint calculation unit further includes a spatial position constraint calculation unit and a grammatical perception contrast constraint calculation unit;
[0140] The spatial position constraint calculation unit includes a spatial perception constraint calculation unit and a background constraint calculation unit, wherein:
[0141] The spatial perception constraint calculation unit is used to generate bounding boxes of the spaces corresponding to multiple subjects in the video based on the Wensheng video input data, and calculate the loss function of the spatial perception constraint as follows:
[0142]
[0143] Among them, Λ fg is the loss function of spatial perception constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated;
[0144] The loss function of the background constraint calculated by the background constraint calculation unit is as follows:
[0145]
[0146] Among them, Λ bg is the loss function of the background constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated;
[0147] The loss function of the spatial position constraint calculated by the spatial position constraint calculation unit is as follows:
[0148] Λ sp =λ fg Λ fg +λ bg Λ bg
[0149] Among them, Λ sp is the loss function of spatial position constraint, λ fg is the loss function of spatial perception constraint Λ fg The weight coefficient, λ bg is the loss function Λ for background constraints bg The weight coefficient of .
[0150] The grammar-aware contrast constraint calculation unit further includes:
[0151] The sample division unit is used to divide the text description S in the video input data into * The word pair consisting of the i-th word (noun) representing the subject name and the j-th word (verb) describing the subject action As positive samples of each other, the text representation S in the video input data *Other word sets in U i As negative samples;
[0152] A forward loss function calculation unit is used to calculate the forward loss function of the grammatical perception contrast constraint to minimize the s in the cross attention map. i and s j The distance between the mappings is expressed as:
[0153]
[0154] in, is the forward loss function of the grammatically aware contrastive constraint, f dist (·) is the distance function calculation between the cross attention maps, and They are the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word and the j-th word in , where F is the number of image frames in the video to be generated;
[0155] Negative loss function calculation unit, used to calculate the negative loss function of the grammatical perception contrast constraint to encourage the cross attention map of the subject and its action word pair to stay away from S * Other word sets in U m The cross attention map of the words in is expressed as:
[0156]
[0157] in, represents the negative loss function of the grammatical-aware contrast constraint, f dist (·) represents the distance function calculation between cross-attention maps, and They are the text description S in the Wensheng video input data * The i-th word and the u-th word in the set of other words U i The cross attention map of the fth frame corresponding to the word in ( ), where F is the number of image frames in the video to be generated;
[0158] The loss function of the grammar-aware contrast constraint is calculated as follows:
[0159]
[0160] Among them, Λ syt Loss function representing the syntax-aware contrastive constraint.
[0161] The above-mentioned video generation system based on the large-scale Vincent video model is an implementation system corresponding to the above-mentioned video generation method based on the large-scale Vincent video model. Its more specific execution steps can refer to the specific embodiment of the above-mentioned video generation method based on the large-scale Vincent video model, and will not be described in detail here.
[0162] Figure 5 and Figure 6 The video generation effects of the video generation method based on the large-scale cultural video model using the spatial and grammatical dual attention guidance of the present invention are shown respectively. Figure 5 This is a schematic diagram comparing the effects of applying the present invention and other existing video generation solutions. Figure 6 This article provides an example of the effect of video generation using the video generation method and system based on a large-scale cultural video model guided by spatial and grammatical dual attention of the present invention, demonstrating the influence of spatial position constraints and grammatical perception constraints on the results.
[0163] like Figure 5 As shown, each group of three images is a screenshot of the generated video. Each column of three images is the result of a common text description driven generation. The colored text at the bottom represents nouns and verbs, and images with the same color are grouped together.
[0164] The first column of images is the generation result of the basic model ModelScope, the second column of images "+LVD" is the result of applying the existing video generation method LVD on ModelScope, and the third column of images "+Ours" is the video generation result of applying the video generation method based on the large-scale cultural video model with dual spatial and grammatical attention guidance proposed in this invention on ModelScope.
[0165] exist Figure 6 In the “w / oΛ sp ” represents the result of using the video generation method proposed by the present invention without imposing spatial position constraints, “w / oΛ bg ” represents the result of using the video generation method proposed in this invention without imposing additional background constraints, “w / oΛ syt ” represents the results of using the video generation method proposed in this invention without imposing grammatical perceptual contrast constraints.
[0166] The above-mentioned video generation system based on the large-scale Vincent video model is an implementation system corresponding to the above-mentioned video generation method based on the large-scale Vincent video model. Its more specific embodiments can refer to the specific embodiments of the above-mentioned video generation method based on the large-scale Vincent video model, and will not be described in detail here.
[0167] like Figure 4As shown, the present invention also provides an electronic device, the electronic device comprising:
[0168] at least one processor; and,
[0169] a memory communicatively connected to at least one processor; wherein,
[0170] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the steps in the aforementioned video generation method based on a large-scale cultural video model.
[0171] Those skilled in the art will appreciate that the structure shown in the figure does not limit the electronic device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0172] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering the various components. Preferably, the power source may be logically connected to the at least one processor 10 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may further include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0173] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0174] Optionally, the electronic device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0175] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0176] The video generation program 12 based on the large-scale Wensheng video model stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, the following steps can be implemented:
[0177] S110: randomly generating a latent space noise map sequence based on a preset large-scale Wensheng video model;
[0178] S120: Denoising the latent space noise map sequence based on the Vincent video input data, and optimizing the latent space noise map sequence within a preset time step in the denoising process using a spatial position constraint and a grammatical-aware contrast constraint; wherein the Vincent video input data includes a text description of the video to be generated, the text description including multiple subjects in the video to be generated and their corresponding actions and background, and bounding boxes of spaces corresponding to the multiple subjects in the video to be generated;
[0179] S130: Decode the latent space noise map sequence obtained in the last step of the denoising process to generate a video.
[0180] The optimization processing of the latent space noise map sequence includes:
[0181] S121: Inputting the latent space noise map sequence of the current step and the text description of the video to be generated into a preset noise prediction network to perform noise prediction;
[0182] S122: Extracting a cross-attention graph sequence generated in the noise prediction network;
[0183] S123: Calculating spatial position constraints and grammatical perception contrast constraints using the cross attention map sequence;
[0184] S124: Calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint with respect to the latent space noise map sequence through back propagation, and update the latent space noise map sequence using the gradients.
[0185] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figure 4 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0186] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0187] The video generation method and system based on a large-scale Vincent video model according to the present invention have been described above by way of example with reference to the accompanying drawings. However, those skilled in the art will appreciate that various improvements may be made to the video generation method and system based on a large-scale Vincent video model according to the present invention without departing from the scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the contents of the appended claims.
Claims
1. A video generation method based on a large-scale Wensheng video model, characterized in that: Video generation for multiple subjects includes the following steps: S110: randomly generating a latent space noise map sequence based on a preset large-scale Wensheng video model; S120: Denoising the latent space noise map sequence based on the Vincent video input data, and optimizing the latent space noise map sequence within a preset time step in the denoising process using a spatial position constraint and a grammatical-aware contrast constraint; wherein the Vincent video input data includes a text description of the video to be generated, the text description including multiple subjects in the video to be generated and their corresponding actions and background, and bounding boxes of spaces corresponding to the multiple subjects in the video to be generated; S130: decoding the latent space noise map sequence obtained in the last step of the denoising process to generate a video; The optimization processing of the latent space noise map sequence includes: S121: Inputting the latent space noise map sequence of the current step and the text description of the video to be generated into a preset noise prediction network to perform noise prediction; S122: Extracting a cross-attention graph sequence generated in the noise prediction network; S123: Calculating spatial position constraints and grammatical perception contrast constraints using the cross attention map sequence; S124: Calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint with respect to the latent space noise map sequence through back propagation, and update the latent space noise map sequence using the gradients.
2. The video generation method based on the large-scale Wensheng video model according to claim 1, characterized in that: The spatial position constraints calculated using the cross-attention map sequence include spatial perception constraints and background constraints; The spatial perception constraint is calculated based on the bounding boxes of the spaces corresponding to the multiple subjects in the generated video in the Wensheng video input data. The loss function of the spatial perception constraint is as follows: Among them, Λ fg is the loss function of spatial perception constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated; The loss function of the background constraint is as follows: Among them, Λ bg is the loss function of the background constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated; The loss function of the spatial position constraint is as follows: L sp =λ fg L fg +λ bg L bg Among them, Λ sp is the loss function of spatial position constraint, λ fg is the loss function of spatial perception constraint Λ fg The weight coefficient, λ bg is the loss function Λ for background constraints bg The weight coefficient of .
3. The video generation method based on the large-scale Wensheng video model according to claim 1, characterized in that: The method of calculating the grammatical-aware contrast constraint using a sequence of cross-attention maps includes: The text description S in the video input data * The word pair consisting of the i-th word representing the subject name and the j-th word describing the subject action As positive samples of each other, the text representation S in the video input data * Other word sets in U i As negative samples; The forward loss function of the grammatical-aware contrast constraint is used to minimize the s in the cross-attention map i and s j The distance between the mappings is expressed as: in, is the forward loss function of the grammatically aware contrastive constraint, f dist (·) is the distance function calculation between the cross attention maps, and They are the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word and the j-th word in , where F is the number of image frames in the video to be generated; The negative loss function of the grammatical-aware contrast constraint is used to encourage the cross-attention map of the subject and its action word pair to stay away from S * The cross attention map of words in the other word set U in is expressed as: in, represents the negative loss function of the grammatical-aware contrast constraint, f dist (·) represents the distance function calculation between cross-attention maps, and They are the text description S in the Wensheng video input data * The f-th frame cross attention map corresponding to the i-th word and the u-th word in , where F is the number of image frames in the video to be generated; The loss function of the grammar-aware contrast constraint is calculated as follows: Among them, Λ syt Loss function representing syntax-aware contrastive constraints.
4. The video generation method based on the large-scale Wensheng video model according to claim 3, characterized in that: The updating of the latent space noise map sequence by using the gradient includes: In the early stage of denoising, the gradient of the cross-attention constraint based on the spatial position is subtracted from each latent space noise map in the latent space noise map sequence at a preset learning rate; In the later stage of denoising, the gradient of the grammatical-aware contrast constraint is subtracted from each latent space noise map in the latent space noise map sequence at a preset learning rate.
5. The video generation method based on the large-scale Wensheng video model according to claim 4, characterized in that: Calculate the gradient to update z at each time step t , as shown below: Among them, α represents the learning rate of the optimization process, λ * represents the weight of the constraint, To calculate the gradient of the latent space noise map sequence obtained by back propagation, t is the current time step and t0 is the number of control update steps.
6. The video generation method based on the large-scale Wensheng video model according to claim 5, characterized in that: The steps S121 to S124 are looped until the maximum number of loop steps is reached.
7. The video generation method based on the large-scale Wensheng video model according to claim 6, characterized in that: The noise prediction network is a U-Net network, which includes three parts: a decoder, an encoder, and a bottleneck layer; wherein, The encoder is used to extract features by reducing the feature map through convolution and pooling; The decoder is symmetrical with the encoder, and performs feature mapping cascade with the encoder after upsampling to retain features and then performs convolution processing; The bottleneck layer includes two 3×3 convolutional layers and one 1×1 convolutional layer.
8. A video generation system based on a large-scale Vincent video model, generating videos for multiple subjects based on the video generation method based on a large-scale Vincent video model according to any one of claims 1 to 7, comprising: A latent space noise map sequence acquisition unit, configured to randomly generate a latent space noise map sequence based on a preset large-scale Wensheng video model; a denoising unit configured to perform denoising processing on the latent space noise map sequence based on Vincent video input data, wherein the Vincent video input data includes a text description of the video to be generated, the text description including a plurality of subjects in the video to be generated and their corresponding actions and background, and bounding boxes of spaces corresponding to the plurality of subjects in the video to be generated; an optimization unit, configured to optimize the latent space noise map sequence within a preset time step in the denoising process using spatial position constraints and grammatical perception contrast constraints; A decoding unit, configured to decode the latent space noise map sequence obtained in the last step of the denoising process to generate a video; Wherein, the optimization unit includes: A noise prediction unit is used to input the latent space noise map sequence of the current step and the text description of the video to be generated into a preset noise prediction network to perform noise prediction; a cross-attention extraction unit, configured to extract a cross-attention graph sequence generated in the noise prediction network; an attention adjustment constraint calculation unit, configured to calculate spatial position constraints and grammatical perception contrast constraints using the cross-attention map sequence; An updating unit is configured to calculate the gradients of the spatial position constraint and the grammatical perception contrast constraint on the latent space noise map sequence through back propagation, and update the latent space noise map sequence using the gradients.
9. The video generation system based on a large-scale cultural video model according to claim 8, characterized in that: The attention adjustment constraint calculation unit further includes a spatial position constraint calculation unit and a grammatical perception contrast constraint calculation unit; wherein, The spatial position constraint calculation unit includes a spatial perception constraint calculation unit and a background constraint calculation unit, wherein: The spatial perception constraint calculation unit is used to generate bounding boxes of the spaces corresponding to multiple subjects in the video based on the Wensheng video input data, and calculate the loss function of the spatial perception constraint as follows: Among them, Λ fg is the loss function of spatial perception constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated; The loss function of the background constraint calculated by the background constraint calculation unit is as follows: Among them, Λ bg is the loss function of the background constraint, is the text description S in the Wensheng video input data * The spatial bounding box of the fth frame corresponding to the i-th word, is the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word, where F is the number of image frames in the video to be generated; The loss function of the spatial position constraint calculated by the spatial position constraint calculation unit is as follows: L sp =λ fg L fg +λ bg L bg Among them, Λ sp is the loss function of spatial position constraint, λ fg is the loss function of spatial perception constraint Λ fg The weight coefficient, λ bg is the loss function Λ for background constraints bg The weight coefficient of The grammar-aware contrast constraint calculation unit includes a grammar-aware contrast constraint positive loss function calculation unit and a grammar-aware contrast constraint negative loss function calculation unit, wherein: The grammatical-aware contrast constraint forward loss function calculation unit is used to calculate the grammatical-aware contrast constraint forward loss function as follows: in, is the forward loss function of the grammatically aware contrastive constraint, f dist (·) is the distance function calculation between the cross attention maps, and They are the text description S in the Wensheng video input data * The cross attention map of the f-th frame corresponding to the i-th word and the j-th word in , where F is the number of image frames in the video to be generated; The grammatical-aware contrast constraint negative loss function calculation unit is used to calculate the grammatical-aware contrast constraint negative loss function as follows: in, represents the negative loss function of the grammatical-aware contrast constraint, U i is the set of other words, f dist (·) represents the distance function calculation between cross-attention maps, and They are the text description S in the Wensheng video input data * The f-th frame cross attention map corresponding to the i-th word and the u-th word in , where F is the number of image frames in the video to be generated; The grammatical-aware contrast constraint loss function calculated by the grammatical-aware contrast constraint calculation unit is as follows: Among them, Λ syt Loss function representing syntax-aware contrastive constraints.
10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the steps in the video generation method based on a large-scale Wensheng video model as described in any one of claims 1 to 7.