Text-based Motion Video Generation Method, Device, Storage Medium and Equipment
By introducing the LCM Lora model, ControlNet model and video enhancement processing model into the AnimateDiff model, the problems of slow video generation speed, discontinuous motion and low clarity are solved, and efficient, accurate and high-definition motion video generation is achieved.
Patent Information
- Application Number
- CN202411206269.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-08-30
AI Technical Summary
The AnimateDiff model generates videos slowly, deformed and discontinuous, and has low video clarity.
The LCM Lora model is introduced to accelerate inference, control action generation using the ControlNet model, and high-definition repair is performed through video enhancement processing models.
Significantly improves video generation speed, ensures consistency and controllability of actions, and significantly improves video clarity.
Smart Images

Figure CN119450164B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method, device, storage medium, and equipment for generating motion videos based on text. Background Art
[0002] Animatediff is an effective framework for extending text-to-image models to animation generators without the need for specific model adaptation. As long as motion priors are learned from large video datasets, AnimateDiff can be inserted into personalized text-to-image models, compatible with existing text-to-image models, or with large models fine-tuned by oneself. However, there are problems with the AnimateDiff model, such as slow video generation speed, distorted and discontinuous actions, and low clarity of the generated videos. Summary of the Invention
[0003] This application provides a method, device, storage medium, and equipment for generating motion videos based on text, which is used to solve the problems of the AnimateDiff model, such as slow video generation speed, distorted and discontinuous actions, and low clarity of the generated videos. The technical solutions are as follows:
[0004] According to the first aspect of this application, a method for generating a motion video based on text is provided. The method includes:
[0005] Obtain a positive prompt and a negative prompt for the motion video to be generated;
[0006] Perform text encoding on the positive prompt and the negative prompt to obtain text features;
[0007] Insert the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model. The Animatediff model is used to endow motion capabilities, and the LCM Lora model is used to accelerate inference;
[0008] Use the video generation model to process the text features to obtain a motion video;
[0009] Use a video enhancement processing model to perform video enhancement processing on the motion video to obtain a high-definition motion video. The video enhancement processing model is used to perform video enhancement.
[0010] In a possible implementation,
[0011] The method further includes: obtaining an action reference video of a person;
[0012] Inserting the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model includes: inserting the Animatediff model, the LCM Lora model, and the ControlNet model into the Stable-Diffusion model to obtain the video generation model, where the ControlNet module is used to provide a reference action;
[0013] Processing the text feature using the video generation model to obtain a motion video includes: processing the text feature and the action reference video using the video generation model to obtain a motion video.
[0014] In a possible implementation, inserting the Animatediff model, the LCM Lora model, and the ControlNet model into the Stable-Diffusion model to obtain the video generation model includes:
[0015] Connecting each module in the LCM Lora model in parallel with the corresponding module in the Stable-Diffusion model to obtain a plurality of parallel modules;
[0016] Adding some modules in the ControlNet model to the corresponding parallel modules to obtain parallel modules;
[0017] Inserting each module in the Animatediff model after each parallel module to obtain a video generation model.
[0018] In a possible implementation, processing the text feature and the action reference video using the video generation model to obtain a motion video includes:
[0019] Processing the text feature using the LCM Lora module and the Stable-Diffusion module in the first parallel module respectively to obtain two first processing results; fusing the two first processing results into a second processing result; processing the action reference video using the j-th ControlNet module, adding the obtained third processing result to the second processing result to obtain a fourth processing result; processing the fourth processing result using the first Animatediff module to obtain a first feature map, where j≥1;
[0020] Use the LCM Lora module and the Stable-Diffusion module in the i-th parallel module to process the first feature map output by the (i - 1)-th Animatediff module respectively, obtaining two first processing results; fuse the two first processing results into a second processing result; use the (j + i - 1)-th ControlNet module to process the third processing result output by the (j + i - 2)-th ControlNet module, and add the obtained third processing result to the second processing result to obtain a fourth processing result; use the i-th Animatediff module to process the fourth processing result to obtain a first feature map, where i ≥ 2;
[0021] Combine the first feature maps output by the last Animatediff module multiple times to form a motion video.
[0022] In a possible implementation manner, the using the video enhancement processing model to perform video enhancement processing on the motion video to obtain a high-definition motion video includes:
[0023] When the video enhancement processing model includes multiple video enhancement modules and a dual regulator, use each video enhancement module to process the motion video to obtain each second feature map; use the dual regulator to adjust each second feature map and the motion video to obtain a high-definition motion video.
[0024] In a possible implementation manner, when the first video enhancement module includes a maximum processing module, an edge prior processing module, an AIE-BAN module, and a three-dimensional convolutional layer, and other video enhancement modules include an AIE-BAN module and a three-dimensional convolutional layer, the using each video enhancement module to process the motion video to obtain each second feature map includes:
[0025] Use the maximum processing module to perform maximum processing on the motion video to obtain a first fifth processing result; use the edge prior processing module to perform edge prior processing on the motion video to obtain a first sixth processing result; use the AIE-BAN module in the first video enhancement module to process the first fifth processing result, the first sixth processing result, and the motion video to obtain a first seventh processing result; use the three-dimensional convolutional layer in the first video enhancement module to perform three-dimensional convolution on the seventh processing result to obtain a first second feature map;
[0026] Perform tensor shape matching on the first fifth processing result to obtain the k-th fifth processing result; perform tensor shape matching on the first sixth processing result to obtain the k-th sixth processing result; use the AIE-BAN module in the k-th video enhancement module to process the k-th fifth processing result, the k-th sixth processing result, and the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module to obtain the k-th seventh processing result; use the three-dimensional convolutional layer in the k-th video enhancement module to perform three-dimensional convolution on the k-th seventh processing result to obtain the k-th second feature map, where k ≥ 2.
[0027] In a possible implementation manner, when the AIE-BAN module includes a convolutional activation unit, a dual regulator, and a sampling unit,
[0028] The step of using the AIE-BAN module in the first video enhancement module to process the first fifth processing result, the first sixth processing result, and the motion video to obtain the first seventh processing result includes: using the convolutional activation unit in the first AIE-BAN module to process and add the motion video and the first fifth processing result respectively to obtain the first eighth processing result; using the dual regulator in the first AIE-BAN module to process the motion video and the first eighth processing result respectively to obtain the first ninth processing result; using the convolutional activation unit to process the first sixth processing result and the first ninth processing result respectively to obtain the first tenth processing result; using the sampling unit in the first AIE-BAN module to process the first tenth processing result to obtain the first seventh processing result;
[0029] The step of using the AIE-BAN module in the k-th video enhancement module to process the k-th fifth processing result, the k-th sixth processing result, and the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module to obtain the k-th seventh processing result includes: using the convolutional activation unit in the k-th AIE-BAN module to process and add the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module and the k-th fifth processing result respectively to obtain the k-th eighth processing result; using the dual regulator in the k-th AIE-BAN module to process the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module and the k-th eighth processing result respectively to obtain the k-th ninth processing result; using the convolutional activation unit to process the k-th sixth processing result and the k-th ninth processing result respectively to obtain the k-th tenth processing result; using the sampling unit in the k-th AIE-BAN module to process the k-th tenth processing result to obtain the k-th seventh processing result.
[0030] According to the second aspect of the present application, there is provided a text-based motion video generation device, which includes:
[0031] An acquisition module, configured to acquire a positive prompt and a negative prompt for the motion video to be generated;
[0032] An encoding module, configured to perform text encoding on the positive prompt and the negative prompt to obtain text features;
[0033] An insertion module, configured to insert the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model, where the Animatediff model is used to endow motion capabilities, and the LCM Lora model is used to accelerate inference;
[0034] A generation module, configured to use the video generation model to process the text features to obtain a motion video;
[0035] The generation module is further configured to perform video enhancement processing on the motion video by using a video enhancement processing model to obtain a high-definition motion video, where the video enhancement processing model is used to perform video enhancement.
[0036] According to the third aspect of the present application, there is provided a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned text-based motion video generation method.
[0037] According to the fourth aspect of the present application, there is provided a computer device, which includes the above-mentioned text-based motion video generation device.
[0038] The beneficial effects of the technical solution provided by the present application at least include:
[0039] The LCM-LoRA model is a lightweight module. By introducing the LCM LoRA module to accelerate the inference of the Stable Diffusion model, the inference process of the model can be optimized, the inference speed is significantly improved, the video generation time is shortened, the video generation efficiency is greatly improved, and the consumption of computing resources is reduced.
[0040] The ControlNet model can capture and encode the action information in the action reference video, and transmit the action information to the Stable Diffusion model, so that the actions in the generated video are consistent with the actions in the action reference video, realizing precise control, thereby ensuring that the actions in the generated video have high consistency and controllability.
[0041] By introducing a video enhancement processing model, frame-by-frame high-definition restoration is performed on the generated motion video, making the generated video clearer and having higher visual quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0043] Figure 1 is a schematic diagram of the network structure of the video generation model;
[0044] Figure 2 is a schematic diagram of the network structure of the video enhancement processing model;
[0045] Figure 3 is a schematic diagram of the network structure of the AIE-BAN model;
[0046] Figure 4 is a flowchart of the method for generating a text-based motion video;
[0047] Figure 5 is a schematic diagram of the method for generating a text-based motion video;
[0048] Figure 6 is a flowchart of the method for generating a text-based motion video;
[0049] Figure 7 is a block diagram of the structure of the device for generating a text-based motion video. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0051] The AnimateDiff model endows the text-to-image model with animation capabilities. The core of the AnimateDiff model lies in inserting a newly initialized motion modeling module into the text-to-image model and training it on a video dataset to extract reasonable motion prior knowledge. Once the training is completed, simply injecting this motion model module can transform the text-to-image model into a text-driven video model, which can generate diverse and personalized dynamic images, covering the fields of anime pictures and realistic photos, while maintaining the domain characteristics and diversity of its output. However, the AnimateDiff model has problems such as slow video generation speed, deformed and discontinuous actions, and low clarity of the generated videos.
[0052] Regarding the problem that the Animatediff model has a slow speed in generating videos, the improvement method is to introduce the Latent Consistency Model (LCM) and the Low-Rank Adaptation (Lora) model to accelerate inference. The LCM Lora model is a general acceleration module without training that can be directly inserted into various StableDiffusion models, supporting resource savings brought by fast inference technology with the fewest steps. Generally speaking, the training of the latent consistency model is carried out by a single-stage guided distillation method. This method uses the pre-trained autoencoder latent space to distill the guided diffusion model into the LCM, so as to ensure that the generated samples follow the trajectory of generating high-quality images.
[0053] Regarding the problem that the Animatediff model has distorted and discontinuous actions, the improvement method is to introduce the ControlNet model to control action generation. Optionally, the actions of the characters in the action reference video can be used as conditions through the ControlNet module to control the batch generation of images with continuous specified actions. The ControlNet model can capture and encode the action information in the action reference video, and transfer the action information to the Stable Diffusion model, so that the actions of the generated video are consistent with the actions in the action reference video, achieving precise control, thereby ensuring that the actions of the generated video have high consistency and controllability.
[0054] Regarding the problem that the videos generated by the Animatediff model have low clarity, the improvement method is to introduce a video enhancement processing model, and use the video enhancement processing model to perform video enhancement processing on the motion video to obtain a high-definition motion video, thereby improving the clarity of the video.
[0055] If it is necessary to use the action reference video to guide the action, the Animatediff model, the LCM Lora model, and the ControlNet model need to be inserted into the Stable-Diffusion model to obtain the first video generation model; if it is not necessary to use the action reference video to guide the action, the Animatediff model and the LCM Lora model need to be inserted into the Stable-Diffusion model to obtain the second video generation model. The model structures of these two video generation models are described below.
[0056] (1) The first video generation model:
[0057] Insert the Animatediff model, LCM Lora model, and ControlNet model into the Stable-Diffusion model to obtain a video generation model, which may include: connecting each module in the LCM Lora model in parallel with the corresponding module in the Stable-Diffusion model to obtain multiple parallel modules; adding some modules in the ControlNet model to the corresponding parallel modules to obtain parallel modules; inserting each module in the Animatediff model after each parallel module to obtain a video generation model.
[0058] As Figure 1 shown, the green hexagon represents the Stable-Diffusion model, the gray square represents the Stable-Diffusion module, the large blue square above represents the Animatediff model, the small blue square in the green hexagon represents the Animatediff module, the large orange square above represents the LCM Lora model, the small orange square in the green hexagon represents the LCM Lora module, the pink hexagon represents the ControlNet model, and the pink square represents the ControlNet module. Figure 1 In, the LCM Lora module and the Stable-Diffusion module form parallel modules, the ControlNet module is added to the parallel modules to form parallel modules, and the Animatediff module is located after each parallel module to combine into a video generation model.
[0059] (2) The second video generation model:
[0060] Insert the Animatediff model and LCM Lora model into the Stable-Diffusion model to obtain a video generation model, which may include: connecting each LCM Lora module in the LCM Lora model in parallel with the corresponding Stable-Diffusion module in the Stable-Diffusion model to obtain multiple parallel modules; inserting each Animatediff module in the Animatediff model after each parallel module to obtain a video generation model.
[0061] As Figure 1 shown, after removing the ControlNet model in the dashed box, all the modules in the green hexagon form a video generation model.
[0062] The model structure of the video enhancement processing model will be described below.
[0063] As Figure 2As shown, the video enhancement processing model includes multiple video enhancement modules and a dual regulator ( Figure 2 not shown in the figure). The first video enhancement module includes a maximum processing module, an edge prior processing module, an AIE-BAN module, and a three-dimensional convolutional layer. The other video enhancement modules include an AIE-BAN module and a three-dimensional convolutional layer. The three-dimensional convolutional layer includes a three-dimensional convolutional module and a non-linear activation function. Figure 2 In the figure, the green hexagon represents the video enhancement processing model. The first green square represents the combination of the maximum processing module, the edge prior processing module, and the AIE-BAN module. The other green squares represent the AIE-BAN module. The pink squares represent the three-dimensional convolutional layer.
[0064] After obtaining the weight parameters of the AIE-BAN module, a video can be used to train the video enhancement processing model. During the training process, it is necessary to fix the weight parameters of the AIE-BAN module and train the weight parameters of the three-dimensional convolutional layer. Specifically, the video enhancement processing model can be used to process the input video to obtain an output video. The mean-square error (MSE) and Structural Similarity (SSIM) are used to calculate the loss value between the input video and the output video, and the weight parameters of the three-dimensional convolutional layer are updated according to the loss value.
[0065] As Figure 3 shown, Figure 3 in the figure, the orange square IL represents the feature map output by the maximum (max) processing module, and the orange square CI represents the feature map output by the edge prior processing (Edge PriorModel, EPM) module. The light pink square and the dark pink square both represent the AIE-BAN module.
[0066] The AIE-BAN module includes a convolutional activation unit, a dual regulator, and a sampling unit. The convolutional activation unit includes a convolution operation + activation function, represented by a gray square; the dual regulator includes a brightener and a darkener, both represented by a yellow square; the sampling unit includes downsampling or upsampling, both represented by a blue square. When the sampling unit in the AIE-BAN module is a downsampling unit, the light pink square represents the AIE-BAN module + downsampling; when the sampling unit in the AIE-BAN module is an upsampling unit, the dark pink square represents the AIE-BAN module + upsampling.
[0067] As Figure 2 shown, any one or more of blurring processing, noise processing, and brightness processing can be performed on the original video to obtain multiple degraded videos. The original video and the corresponding degraded videos are combined into a set of training samples, and the video enhancement processing model is trained using multiple sets of training samples.
[0068] AsFigure 4 As shown, it shows a method flow chart of a text-based motion video generation method provided by an embodiment of the present application. This text-based motion video generation method can be applied to a computer device. This text-based motion video generation method may include:
[0069] Step 401, obtain a positive prompt word and a negative prompt word for the motion video to be generated.
[0070] The positive prompt word is used to describe the elements expected to appear in the video. For example, a red-haired girl dancing, etc.
[0071] The negative prompt word is used to describe the elements not expected to appear in the video. For example, of poor quality, blurred, etc.
[0072] Step 402, perform text encoding on the positive prompt word and the negative prompt word to obtain text features.
[0073] In this embodiment, a text encoder can be used to encode the positive prompt word and the negative prompt word respectively to obtain two text features.
[0074] Step 403, insert the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model. The Animatediff model is used to endow the motion ability, and the LCM Lora model is used to accelerate the inference.
[0075] After inserting the Animatediff model and the LCM Lora model into the Stable-Diffusion model, a video generation model can be obtained. The structure of this video generation model is as Figure 1 shown.
[0076] Step 404, use the video generation model to process the text features to obtain a motion video.
[0077] Specifically, using the video generation model to process the text features to obtain a motion video may include:
[0078] (1) Use the LCM Lora module and the Stable-Diffusion module in the first parallel module to process the text features respectively to obtain two first processing results; fuse the two first processing results into a second processing result; use the first Animatediff module to process the second processing result to obtain a third feature map.
[0079] (2) Use the LCM Lora module and the Stable-Diffusion module in the i-th parallel module to process the third feature map output by the (i-1)-th Animatediff module respectively, obtaining two first processing results; fuse the two first processing results into a second processing result; use the i-th Animatediff module to process the second processing result to obtain a third feature map, where i≥2.
[0080] (3) Combine the third feature maps output multiple times by the last Animatediff module into a motion video.
[0081] Simply put, the input of the first parallel module is text features, the input of the i-th (i≥2) parallel module is the third feature map output by the (i-1)-th parallel module, the last parallel module outputs the third feature map multiple times, and all the third feature maps are combined into a motion video.
[0082] Step 405, use a video enhancement processing model to perform video enhancement processing on the motion video to obtain a high-definition motion video. The video enhancement processing model is used for video enhancement.
[0083] The video enhancement processing model can perform any one or more of blurring processing, noise processing, and brightness processing on the motion video to obtain a high-definition motion video.
[0084] As Figure 5 shown, insert the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model; encode the positive prompt and the negative prompt and input them into the video generation model; use the video generation model to process the text features to obtain a motion video; use the video enhancement processing model to process the motion video to obtain a high-definition motion video.
[0085] In summary, for the text-based motion video generation method provided in the embodiments of the present application, the LCM-LoRA model is a lightweight module. By introducing the LCM LoRA module to accelerate the inference of the Stable Diffusion model, the inference process of the model can be optimized, the inference speed is significantly improved, the video generation time is shortened, the video generation efficiency is greatly improved, and the consumption of computing resources is reduced.
[0086] By introducing a video enhancement processing model to perform frame-by-frame high-definition repair on the generated motion video, the generated video is made clearer and has higher visual quality.
[0087] As Figure 6As shown, it shows a flowchart of a text-based motion video generation method provided by an embodiment of the present application. This text-based motion video generation method can be applied to a computer device. This text-based motion video generation method may include:
[0088] Step 601, obtain a positive prompt and a negative prompt for the motion video to be generated, and obtain an action reference video of a person.
[0089] The positive prompt is used to describe the elements expected to appear in the video. For example, a red-haired girl dancing, etc.
[0090] The negative prompt is used to describe the elements not expected to appear in the video. For example, of poor quality, blurred, etc.
[0091] The action reference video can be a video containing reference actions. For example, a person's dancing video.
[0092] Step 602, perform text encoding on the positive prompt and the negative prompt to obtain text features.
[0093] In this embodiment, the text encoder can be used to encode the positive prompt and the negative prompt respectively to obtain two text features.
[0094] Step 603, insert the Animatediff model, the LCM Lora model, and the ControlNet model into the Stable-Diffusion model to obtain a video generation model. The ControlNet module is used to provide reference actions.
[0095] After inserting the Animatediff model, the LCM Lora model, and the ControlNet model into the Stable-Diffusion model, a video generation model can be obtained. The structure of this video generation model is as Figure 1 shown.
[0096] Step 604, use the video generation model to process the text features and the action reference video to obtain a motion video.
[0097] Specifically, using the video generation model to process the text features and the action reference video to obtain a motion video may include:
[0098] (1) Use the LCM Lora module and the Stable-Diffusion module in the first parallel module to process the text features respectively, obtaining two first processing results; fuse the two first processing results into a second processing result; use the j-th ControlNet module to process the action reference video, and add the obtained third processing result to the second processing result to obtain a fourth processing result; use the first Animatediff module to process the fourth processing result to obtain a first feature map, where j≥1.
[0099] The value of j can be set according to actual needs. Taking Figure 1 as an example, start inserting from the central ControlNet module in the ControlNet model. Assuming there are 5 ControlNet modules in the ControlNet model, the value of j is 3; assuming there are 7 ControlNet modules in the ControlNet model, the value of j is 4, and so on.
[0100] (2) Use the LCM Lora module and the Stable-Diffusion module in the i-th parallel module to process the first feature map output by the (i - 1)-th Animatediff module respectively, obtaining two first processing results; fuse the two first processing results into a second processing result; use the (j + i - 1)-th ControlNet module to process the third processing result output by the (j + i - 2)-th ControlNet module, and add the obtained third processing result to the second processing result to obtain a fourth processing result; use the i-th Animatediff module to process the fourth processing result to obtain a first feature map, where i≥2.
[0101] (3) Combine the first feature maps output multiple times by the last Animatediff module into a motion video.
[0102] Simply put, the input of the first parallel module is text features, the input of the i-th (i≥2) parallel module is the first feature map output by the (i - 1)-th parallel module, and the last parallel module outputs the first feature map multiple times. Combine all the first feature maps into a motion video.
[0103] Step 605, when the video enhancement processing model includes multiple video enhancement modules and a dual regulator, use each video enhancement module to process the motion video to obtain respective second feature maps; use the dual regulator to adjust each second feature map and the motion video to obtain a high-definition motion video.
[0104] Among them, the dual regulator includes a brightener and a darkener. The calculation formula of the brightener is The calculation formula of the darkener is Among them, represents the weight map, I(x) represents the input, and γ represents a constant.
[0105] When the first video enhancement module includes a maximum processing module, an edge prior processing module, an AIE-BAN module, and a three-dimensional convolutional layer, and other video enhancement modules include an AIE-BAN module and a three-dimensional convolutional layer, processing the motion video using each video enhancement module to obtain each second feature map may include:
[0106] (1) Using the maximum processing module to perform maximum processing on the motion video to obtain the first fifth processing result; using the edge prior processing module to perform edge prior processing on the motion video to obtain the first sixth processing result; using the AIE-BAN module in the first video enhancement module to process the first fifth processing result, the first sixth processing result, and the motion video to obtain the first seventh processing result; using the three-dimensional convolutional layer in the first video enhancement module to perform three-dimensional convolution on the first seventh processing result to obtain the first second feature map.
[0107] (2) Performing tensor shape matching on the first fifth processing result to obtain the kth fifth processing result; performing tensor shape matching on the first sixth processing result to obtain the kth sixth processing result; using the AIE-BAN module in the kth video enhancement module to process the kth fifth processing result, the kth sixth processing result, and the (k - 1)th second feature map output by the (k - 1)th video enhancement module to obtain the kth seventh processing result; using the three-dimensional convolutional layer in the kth video enhancement module to perform three-dimensional convolution on the kth seventh processing result to obtain the kth second feature map, where k ≥ 2.
[0108] Tensor matching means adjusting the first fifth processing result and the first sixth processing result to have the same size as the (k - 1)th second feature map, which may remain unchanged or may be downsampled by a factor of 2 or 4 or more.
[0109] Simply put, the input of the first video enhancement module is the motion video, and the input of the kth (k ≥ 2) video enhancement module is the (k - 1)th second feature map output by the (k - 1)th video enhancement module.
[0110] Among them, when the AIE-BAN module includes a convolutional activation unit, a dual regulator, and a sampling unit,
[0111] (1) The AIE-BAN module in the first video enhancement module processes the first fifth processing result, the first sixth processing result, and the motion video to obtain the first seventh processing result, including: adding the results obtained by processing the motion video and the first fifth processing result respectively by the convolutional activation units in the first AIE-BAN module to obtain the first eighth processing result; processing the motion video and the first eighth processing result respectively by the dual regulators in the first AIE-BAN module to obtain the first ninth processing result; processing the first sixth processing result and the first ninth processing result respectively by the convolutional activation units to obtain the first tenth processing result; and processing the first tenth processing result by the sampling unit in the first AIE-BAN module to obtain the first seventh processing result.
[0112] The AIE-BAN module may include multiple convolutional activation units. As Figure 3 shown, in one embodiment, the motion video is processed by 3 convolutional activation units, the fifth processing result is processed by 3 convolutional activation units, and the sixth processing result and the ninth processing result are processed by 3 convolutional activation units.
[0113] (2) The AIE-BAN module in the k-th video enhancement module processes the k-th fifth processing result, the k-th sixth processing result, and the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module to obtain the k-th seventh processing result, including: adding the results obtained by processing the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module and the k-th fifth processing result respectively by the convolutional activation units in the k-th AIE-BAN module to obtain the k-th eighth processing result; processing the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module and the k-th eighth processing result respectively by the dual regulators in the k-th AIE-BAN module to obtain the k-th ninth processing result; processing the k-th sixth processing result and the k-th ninth processing result respectively by the convolutional activation units to obtain the k-th tenth processing result; and processing the k-th tenth processing result by the sampling unit in the k-th AIE-BAN module to obtain the k-th seventh processing result.
[0114] In short, the input of the convolutional activation unit in the first AIE-BAN module is the motion video, and the input of the convolutional activation unit in the k-th (k ≥ 2) AIE-BAN module is the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module.
[0115] As Figure 5As shown in the figure, the Animatediff model, LCM Lora model, and ControlNet model are inserted into the Stable-Diffusion model to obtain a video generation model. The positive prompt and negative prompt are encoded and then input into the video generation model. The video generation model processes the text features to obtain a motion video. The motion video is processed by a video enhancement processing model to obtain a high-definition motion video.
[0116] In summary, for the text-based motion video generation method provided in the embodiments of the present application, the LCM-LoRA model is a lightweight module. By introducing the LCM LoRA module to accelerate the inference of the Stable Diffusion model, the inference process of the model can be optimized, the inference speed is significantly improved, the video generation time is shortened, the video generation efficiency is greatly improved, and the consumption of computing resources is reduced.
[0117] The ControlNet model can capture and encode the action information in the action reference video, and transmit the action information to the Stable Diffusion model, so that the actions in the generated video are consistent with the actions in the action reference video, realizing precise control, thereby ensuring that the actions in the generated video have high consistency and controllability.
[0118] By introducing a video enhancement processing model to perform frame-by-frame high-definition repair on the generated motion video, the generated video is made clearer and has higher visual quality.
[0119] As Figure 7 shown in the figure, it shows the structural block diagram of a text-based motion video generation device provided in an embodiment of the present application. This text-based motion video generation device can be applied to a computer device. The text-based motion video generation device may include:
[0120] An acquisition module 710, configured to acquire the positive prompt and negative prompt for the motion video to be generated;
[0121] An encoding module 720, configured to perform text encoding on the positive prompt and negative prompt to obtain text features;
[0122] An insertion module 730, configured to insert the Animatediff model and LCM Lora model into the Stable-Diffusion model to obtain a video generation model. The Animatediff model is used to endow motion ability, and the LCM Lora model is used to accelerate inference;
[0123] A generation module 740, configured to use the video generation model to process the text features to obtain a motion video;
[0124] The generation module 740 is further configured to perform video enhancement processing on the motion video by using a video enhancement processing model to obtain a high-definition motion video, and the video enhancement processing model is used for video enhancement.
[0125] In an optional embodiment, the acquisition module 710 is further configured to acquire an action reference video of a person;
[0126] The insertion module 730 is further configured to insert the Animatediff model, the LCM Lora model, and the ControlNet model into the Stable-Diffusion model to obtain a video generation model, and the ControlNet module is used to provide reference actions;
[0127] The generation module 740 is further configured to process the text features and the action reference video by using the video generation model to obtain a motion video.
[0128] In an optional embodiment, the insertion module 730 is further configured to include:
[0129] Parallelly connect each module in the LCM Lora model with the corresponding module in the Stable-Diffusion model to obtain a plurality of parallel modules;
[0130] Add some modules in the ControlNet model to the corresponding parallel modules to obtain a parallel module group;
[0131] Insert each module in the Animatediff model after each parallel module group to obtain a video generation model.
[0132] In an optional embodiment, the generation module 740 is further configured to:
[0133] Process the text features by using the LCM Lora module and the Stable-Diffusion module in the first parallel module respectively to obtain two first processing results; fuse the two first processing results into a second processing result; process the action reference video by using the j-th ControlNet module, add the obtained third processing result to the second processing result to obtain a fourth processing result; process the fourth processing result by using the first Animatediff module to obtain a first feature map, where j≥1;
[0134] The LCM Lora module and the Stable-Diffusion module in the $i$-th parallel module are used to process the first feature map output by the $(i - 1)$-th Animatediff module respectively, obtaining two first processing results; the two first processing results are fused into a second processing result; the $(j + i - 1)$-th ControlNet module is used to process the third processing result output by the $(j + i - 2)$-th ControlNet module, and the obtained third processing result is added to the second processing result to obtain a fourth processing result; the $i$-th Animatediff module is used to process the fourth processing result to obtain a first feature map, where $i\geq2$.
[0135] The first feature maps output by the last Animatediff module are combined into a motion video.
[0136] In an optional embodiment, the generating module 740 is further configured to:
[0137] When the video enhancement processing model includes multiple video enhancement modules and a dual regulator, each video enhancement module is used to process the motion video to obtain respective second feature maps; the dual regulator is used to adjust each second feature map and the motion video to obtain a high-definition motion video.
[0138] In an optional embodiment, when the first video enhancement module includes a maximum processing module, an edge prior processing module, an AIE-BAN module, and a three-dimensional convolutional layer, and the other video enhancement modules include an AIE-BAN module and a three-dimensional convolutional layer, the generating module 740 is further configured to:
[0139] The maximum processing module is used to perform maximum processing on the motion video to obtain a first fifth processing result; the edge prior processing module is used to perform edge prior processing on the motion video to obtain a first sixth processing result; the AIE-BAN module in the first video enhancement module is used to process the first fifth processing result, the first sixth processing result, and the motion video to obtain a first seventh processing result; the three-dimensional convolutional layer in the first video enhancement module is used to perform three-dimensional convolution on the first seventh processing result to obtain a first second feature map;
[0140] Perform tensor shape matching on the first fifth processing result to obtain the k-th fifth processing result; perform tensor shape matching on the first sixth processing result to obtain the k-th sixth processing result; use the AIE-BAN module in the k-th video enhancement module to process the k-th fifth processing result, the k-th sixth processing result, and the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module to obtain the k-th seventh processing result; use the 3D convolutional layer in the k-th video enhancement module to perform 3D convolution on the k-th seventh processing result to obtain the k-th second feature map, where k ≥ 2.
[0141] In an optional embodiment, when the AIE-BAN module includes a convolutional activation unit, a dual regulator, and a sampling unit, the generation module 740 is further configured to:
[0142] Use the convolutional activation unit in the first AIE-BAN module to process and add the motion video and the first fifth processing result respectively to obtain the first eighth processing result; use the dual regulator in the first AIE-BAN module to process the motion video and the first eighth processing result respectively to obtain the first ninth processing result; use the convolutional activation unit to process the first sixth processing result and the first ninth processing result respectively to obtain the first tenth processing result; use the sampling unit in the first AIE-BAN module to process the first tenth processing result to obtain the first seventh processing result;
[0143] Use the convolutional activation unit in the k-th AIE-BAN module to process and add the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module and the k-th fifth processing result respectively to obtain the k-th eighth processing result; use the dual regulator in the k-th AIE-BAN module to process the (k - 1)-th second feature map output by the (k - 1)-th video enhancement module and the k-th eighth processing result respectively to obtain the k-th ninth processing result; use the convolutional activation unit to process the k-th sixth processing result and the k-th ninth processing result respectively to obtain the k-th tenth processing result; use the sampling unit in the k-th AIE-BAN module to process the k-th tenth processing result to obtain the k-th seventh processing result.
[0144] In summary, for the text-based motion video generation device provided in the embodiments of the present application, the LCM-LoRA model is a lightweight module. By introducing the LCM LoRA module to accelerate the inference of the Stable Diffusion model, the inference process of the model can be optimized, the inference speed is significantly improved, the video generation time is shortened, the video generation efficiency is greatly improved, and the consumption of computing resources is reduced.
[0145] The ControlNet model can capture and encode the action information in the action reference video, and transfer the action information to the Stable Diffusion model, so that the actions in the generated video are consistent with those in the action reference video, achieving precise control, thereby ensuring that the actions in the generated video have high consistency and controllability.
[0146] By introducing a video enhancement processing model to perform frame-by-frame high-definition repair on the generated motion video, the generated video becomes clearer and has higher visual quality.
[0147] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned text-based motion video generation method.
[0148] An embodiment of the present application provides a computer device, which includes any of the above text-based motion video generation devices.
[0149] It should be noted that: when the above-mentioned text-based motion video generation device provided in the embodiment performs portrait control generation of a specific person, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the text-based motion video generation device is divided into different functional modules to complete all or part of the functions described above. In addition, the above-mentioned text-based motion video generation device provided in the embodiment and the text-based motion video generation method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0150] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0151] The above does not intend to limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.
Claims
1. A method for generating a motion video based on text, characterized in that: The method comprises: Obtaining positive prompt words and negative prompt words for the motion video to be generated; Performing text encoding on the positive prompt words and the negative prompt words to obtain text features; Insert the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model, wherein the Animatediff model is used to impart motion capability, and the LCM Lora model is used to accelerate reasoning; Processing the text features using the video generation model to obtain a motion video; The motion video is subjected to video enhancement processing by using a video enhancement processing model to obtain a high-definition motion video, wherein the video enhancement processing model is used for performing video enhancement.
2. The method for generating a motion video based on text according to claim 1, characterized in that: The method further comprises: obtaining a reference video of the action of the character; The inserting of the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain the video generation model includes: inserting the Animatediff model, the LCM Lora model and the ControlNet model into the Stable-Diffusion model to obtain the video generation model, wherein the ControlNet module is used to provide reference actions; The method of using the video generation model to process the text features to obtain a motion video includes: using the video generation model to process the text features and the action reference video to obtain a motion video.
3. The method for generating a text-based motion video according to claim 2, characterized in that: The method of inserting the Animatediff model, the LCM Lora model and the ControlNet model into the Stable-Diffusion model to obtain the video generation model includes: Connecting each module in the LCM Lora model in parallel with a corresponding module in the Stable-Diffusion model to obtain a plurality of parallel modules; Adding some modules in the ControlNet model to corresponding parallel modules to obtain parallel modules; After inserting each module in the Animatediff model into each parallel module, a video generation model is obtained.
4. The method for generating a text-based motion video according to claim 3, characterized in that: The step of using the video generation model to process the text features and the action reference video to obtain a motion video includes: The text features are processed respectively by the LCM Lora module and the Stable-Diffusion module in the first parallel module to obtain two first processing results; the two first processing results are fused into a second processing result; the action reference video is processed by the jth ControlNet module, and the third processing result obtained is added to the second processing result to obtain a fourth processing result; the fourth processing result is processed by the first Animatediff module to obtain a first feature map, j≥1; The first feature map output by the i-1th Animatediff module is processed by the LCM Lora module and the Stable-Diffusion module in the i-th parallel module to obtain two first processing results; the two first processing results are merged into a second processing result; the third processing result output by the j+i-2th ControlNet module is processed by the j+i-1th ControlNet module, and the third processing result is added to the second processing result to obtain a fourth processing result; the fourth processing result is processed by the i-th Animatediff module to obtain a first feature map, i≥2; Combine the first feature maps output by the last Animatediff module multiple times into a motion video.
5. The method for generating a motion video based on text according to claim 1, characterized in that: The step of performing video enhancement processing on the motion video using the video enhancement processing model to obtain a high-definition motion video includes: When the video enhancement processing model includes multiple video enhancement modules and dual regulators, each video enhancement module is used to process the motion video to obtain each second feature map; the dual regulator is used to adjust each second feature map and the motion video to obtain a high-definition motion video.
6. The method for generating a text-based motion video according to claim 5, characterized in that: When the first video enhancement module includes a maximum value processing module, an edge prior processing module, an AIE-BAN module and a three-dimensional convolution layer, and the other video enhancement modules include an AIE-BAN module and a three-dimensional convolution layer, the motion video is processed by each video enhancement module to obtain each second feature map, including: Using the maximum value processing module to perform maximum value processing on the motion video to obtain a first fifth processing result; using the edge prior processing module to perform edge prior processing on the motion video to obtain a first sixth processing result; using the AIE-BAN module in the first video enhancement module to process the first fifth processing result, the first sixth processing result and the motion video to obtain a first seventh processing result; using the three-dimensional convolution layer in the first video enhancement module to perform three-dimensional convolution on the first seventh processing result to obtain a first second feature map; Perform tensor shape matching on the first fifth processing result to obtain the kth fifth processing result; perform tensor shape matching on the first sixth processing result to obtain the kth sixth processing result; use the AIE-BAN module in the kth video enhancement module to process the kth fifth processing result, the kth sixth processing result and the k-1th second feature map output by the k-1th video enhancement module to obtain the kth seventh processing result; use the three-dimensional convolution layer in the kth video enhancement module to perform three-dimensional convolution on the kth seventh processing result to obtain the kth second feature map, k≥2.
7. The method for generating a motion video based on text according to claim 6, characterized in that: When the AIE-BAN module includes a convolutional activation unit, a dual regulator, and a sampling unit, The using the AIE-BAN module in the first video enhancement module to process the first fifth processing result, the first sixth processing result and the motion video to obtain the first seventh processing result includes: using the convolution activation unit in the first AIE-BAN module to process the motion video and the first fifth processing result respectively and then add them to obtain the first eighth processing result; using the dual regulator in the first AIE-BAN module to process the motion video and the first eighth processing result respectively to obtain the first ninth processing result; using the convolution activation unit to process the first sixth processing result and the first ninth processing result respectively to obtain the first tenth processing result; using the sampling unit in the first AIE-BAN module to process the first tenth processing result to obtain the first seventh processing result; The method of using the AIE-BAN module in the kth video enhancement module to process the kth fifth processing result, the kth sixth processing result and the k-1th second feature map output by the k-1th video enhancement module to obtain the kth seventh processing result includes: using the convolution activation unit in the kth AIE-BAN module to process the k-1th second feature map output by the k-1th video enhancement module and the kth fifth processing result respectively and then add them to obtain the kth eighth processing result; using the dual regulator in the kth AIE-BAN module to process the k-1th second feature map output by the k-1th video enhancement module and the kth eighth processing result respectively to obtain the kth ninth processing result; using the convolution activation unit to process the kth sixth processing result and the kth ninth processing result respectively to obtain the kth tenth processing result; using the sampling unit in the kth AIE-BAN module to process the kth tenth processing result to obtain the kth seventh processing result.
8. A text-based motion video generation device, characterized in that: The device comprises: An acquisition module, used to acquire positive prompt words and negative prompt words of the motion video to be generated; An encoding module, used for performing text encoding on the positive prompt words and the negative prompt words to obtain text features; An insertion module is used to insert the Animatediff model and the LCM Lora model into the Stable-Diffusion model to obtain a video generation model, wherein the Animatediff model is used to impart motion capability, and the LCM Lora model is used to accelerate reasoning; A generation module, used for processing the text features using the video generation model to obtain a motion video; The generation module is also used to perform video enhancement processing on the motion video using a video enhancement processing model to obtain a high-definition motion video, and the video enhancement processing model is used to perform video enhancement.
9. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the text-based motion video generation method according to any one of claims 1 to 7.
10. A computer device, characterized in that: The computer device comprises: the text-based motion video generating device as described in claim 8.
Citation Information
Patent Citations
Anchor video model training method, anchor video generation method and related devices
CN117915164A
Personalized text-to-image generation method and system based on diffusion model
CN118429485A