Text Prompt and Image-Driven Content Generation Method, Device, Medium
By introducing conditional encoding modules and inter-frame consistency encoding and retaining image details in the image driving model, the problem of large differences between the generation results and the original image in the prior art and the inability to control the action amplitude is solved, and a more stable and controllable image driving effect is achieved.
Patent Information
- Application Number
- CN202311759693.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-12-20
AI Technical Summary
The existing image driving technology has problems such as large differences in the generation results from the original image, uncontrollable content of the generated content, and uncontrollable action amplitude, and it is impossible to effectively restore image details and respond to text prompts.
By introducing a conditional encoding module into the image driving model, combining inter-frame consistency encoding, explicitly encoding and retaining the detailed information of a given conditional frame, and by fine-tuning the conditional encoding module and timing module, the response ability and dynamic control of text prompt words are improved.
It realizes better encoding and preservation of image details, improves the stability and controllability of generated videos, and can explicitly control the intensity of the dynamic effect, so that the generated videos are more loyal to the original image content and text prompts.
Smart Images

Figure CN117911584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image driving technology, and more particularly to a content generation method, device, and medium based on text prompts and image driving. Background Art
[0002] Image driving technology refers to the technology of generating a video starting from an image given by the user. The existing image driving technology has the following disadvantages:
[0003] (1) The generated result has a large difference from the original image of the given image. When the model drives the image, the generated video by the existing method has a low similarity with the original image, and even changes the content of the original image, especially some detailed features of the object will be significantly changed.
[0004] (2) The generated content cannot be controlled. When generating a specified dynamic effect with text as a condition, the existing method often fails to respond to the content specified in the generated text and cannot achieve the effect of text control.
[0005] (3) The action amplitude of the generated content cannot be controlled. When generating the dynamic effect described in the text, the existing method lacks a method for controlling the action amplitude and cannot control the action amplitude.
[0006] The image driving technology based on text prompts means that when the user is given an image, the user can generate a video starting from the given image by giving text prompts. The generated video should restore the image content, retain the image details and the features of people and objects; and generate a smooth dynamic effect that conforms to the text prompts according to the text information.
[0007] The existing technical solutions mainly implement image driving based on the text-video generation method (AnimateDiff) of the diffusion model and the control network (ControlNet). In this technical solution, first, the text prompts are encoded by the CLIP text encoder, and then the obtained text encoding is input into the diffusion model to guide the generation process to conform to the text prompts; among them, the diffusion model is obtained by expanding the two-dimensional model (Stable Diffusion) of text-picture in the time dimension, and this expansion enables the model to generate continuous videos without flickering; at the same time, the given picture will be subjected to feature extraction through the control network ControlNet and act on the feature map of the diffusion model, so as to control the generated result of the video to conform to the given image. However, this technology has the following disadvantages:
[0008] (1) Generate a video by combining the AnimateDiff model that generates videos from text, and at the same time use ControlNet to inject image information during the generation process to obtain the effect of image control. Such a technical solution is limited in that during the process of feature extraction of a given image by ControlNet, the detailed information of the image will be lost to some extent, resulting in the inability to restore the details in the given image after image driving. Specifically, although the main content and structure of the input image can be maintained, the features of people and objects and the details of the picture cannot be restored.
[0009] (2) During the training process of the existing solutions, text prompts mainly play a role in controlling the main visual content of the picture, which limits the control of dynamic effects related to actions and events. Specifically, when the text is used as a condition, the response ability to the text condition is poor and it is weakly controlled by the text. When there is a mismatch between the text prompt and the generated image content, the existing methods often can only generate relatively static pictures and lack a response to the text, and cannot make the image move according to the description of the text prompt.
[0010] (3) The inputs of the existing solutions are only text prompts and images, and users cannot explicitly control the degree of change of the picture content. Given an image and the corresponding text prompt, the existing technical solutions can only obtain some random results and cannot explicitly control the intensity of the picture dynamic effect.
[0011] In summary, there is currently a lack of an image-driven content generation method to overcome or partially overcome the aforementioned problems. Summary of the Invention
[0012] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a content generation method based on text prompts and image driving to better encode and retain the details of the given conditional frame.
[0013] The purpose of the present invention can be achieved by the following technical solutions:
[0014] In one aspect of the present invention, there is provided a content generation method based on text prompts and image driving. Based on the given text prompts and the given image, a pre-trained image driving model is used to generate a video. The training process of the image driving model includes the following steps:
[0015] Obtain samples including input text, given conditional frames, target video frame sequences, and inter-frame consistency encoding, wherein the inter-frame consistency encoding is calculated based on the given conditional frames and the target video frame sequences;
[0016] Encode the given conditional frames to obtain image encodings, and based on the image encodings and the inter-frame consistency encoding, obtain conditional frame features through conditional encoding;
[0017] Initialize the noise frame and obtain the noise features through feature extraction;
[0018] Based on the conditional frame features, the noise features, and the input text, obtain the output encoding and perform denoising, which is used as the new noise frame to complete this round of iteration. Repeat this step for multiple iterations;
[0019] Based on the denoised output encoding after multiple iterations, obtain the output video frame, and update the parameters of the image driving model based on the target video frame sequence and the output video frame to complete the training for the sample.
[0020] Among them, conditional encoding is implemented using a conditional encoding module. The conditional encoding module is inserted into the first layer of the text-to-video generation model. The conditional encoding module includes a convolutional layer with a structure of 4×3×3×320. The conditional encoding module uses 320 convolutional kernels with a size of 4×3×3 to perform conditional encoding on the input with 4 channels, and obtains the conditional encoding with the same size as the input and 320 channels.
[0021] As a preferred technical solution, the image driving model includes:
[0022] A conditional encoding module, which is used to obtain conditional frame features based on the image encoding and the inter-frame consistency encoding;
[0023] A raw input module, which is used to obtain noise features based on the noise frame;
[0024] At least one group of Unet modules and a timing module, which are used to obtain the output encoding based on the conditional frame features, the noise features, and the text encoding obtained by encoding the input text.
[0025] As a preferred technical solution, the Unet module is used to process the video frames frame by frame based on the text encoding, and the timing module is used to align the video frames.
[0026] As a preferred technical solution, the raw input module and the Unet module are pre-trained and their parameters are not updated during the training of the image driving model.
[0027] As a preferred technical solution, the conditional frame features are obtained by concatenating the image encoding and the inter-frame consistency encoding on the channel dimension and performing polar conditional encoding.
[0028] As a preferred technical solution, the calculation process of the inter-frame consistency encoding includes:
[0029] Calculate the 1-norm distance between the given conditional frame and each frame in the target video frame sequence in a preset color space, and perform global normalization processing based on the maximum value in the sample set to obtain the inter-frame consistency encoding.
[0030] As a preferred technical solution, for the 1-norm distances below a% or above b% in the sample set, they are respectively replaced with the 1-norm distances at a% and b%.
[0031] As a preferred technical solution, the inter-frame consistency encoding is calculated using the following formula:
[0032]
[0033]
[0034]
[0035] Among them, is the 1-norm distance between the th given conditional frame and the th frame in the target video frame sequence in the HSV color space. This distance measures the difference in the motion amplitude between each frame in the training data and the conditional frame. is the maximum 1-norm distance in the sample set. is the obtained inter-frame consistency encoding. and are the normalization hyperparameters. represents the 1-norm distance corresponding to the 5% statistic in the sample set. represents the 1-norm distance corresponding to the 95% statistic. is defined as.
[0036] Another aspect of the present invention provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the above-mentioned text prompt word and image-driven content generation method.
[0037] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the above-mentioned text prompt word and image-driven content generation method.
[0038] Compared with the prior art, the present invention has at least one of the following advantages:
[0039] (1) Better encoding and retention of details of a given conditional frame: By constructing a conditional encoding module for image-driven tasks, this module can be compatible with existing text-to-image and text-to-video models. Taking the given conditional frame and inter-frame consistency as inputs, it can better encode and retain the details of the conditional frame, retain the effective information of the given conditional frame, and add it to the encoding of the noise by the original input module through encoding. Through such a design, it is possible to explicitly encode and retain the information and details in the conditional frame according to the input inter-frame consistency, and generate video segments that better restore the generated image content.
[0040] (2) Improving the stability and controllability of the generated video: The expansion of the existing training dataset in this application, in addition to the conditional frame and the target conditional frame, also includes the inter-frame consistency between the conditional frame and the target conditional frame. By explicitly encoding the conditional frame information through the conditional encoding model and fine-tuning the conditional encoding model and the temporal model simultaneously, it shows a better response to text prompts and can effectively generate relevant dynamic effects.
[0041] (3) Controllability of the intensity of dynamic effects: Map data with too fast or static actions to a specific input interval, and avoid this input interval during inference to obtain high-quality generation results. At the same time, after training, the intensity of the dynamic effects in the generated video can be controlled by adjusting the value of the input inter-frame consistency. The input of inter-frame consistency can not only allow users to avoid setting extremely small values to avoid low-quality images with overly intense animations, but also allow users to adjust the intensity of the dynamic effects in the generated video within a reasonable numerical range to achieve a more controllable image-driven effect. Description of the Drawings
[0042] Figure 1 It is a schematic flowchart of the content generation method based on text prompts and image driving in the embodiment;
[0043] Figure 2 It is a schematic diagram of driving an image by moderately responding to text prompts in the embodiment;
[0044] Figure 3 It is a schematic diagram of the effect of adjusting the generation of dynamic effects with different degrees of change in the embodiment. Detailed Embodiments
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] Embodiment 1
[0047] In view of the problems existing in the foregoing prior art, in order to accurately restore the structure, style, and details of the input image in the video generated by the driving image, make the video generated by the driving image better controlled by the text, be able to generate dynamic effects that conform to the text prompt, and make the video generated by the driving image more controllable. Specifically, it is to achieve the ability to control the change range of the dynamic effects in the generated video. This embodiment provides a content generation method based on text prompts and image driving.
[0048] See Figure 1 , this embodiment mainly includes the following three aspects.
[0049] (1) Conditional encoding module. This conditional encoding module can be compatible with the existing model for generating videos from text. Inserting this conditional encoding module into the existing model for generating videos from text can obtain an image-driven model. This conditional encoding module can not only retain the response ability of the image-driven model to text prompts, but also retain the high-frequency detail information of the given picture by explicitly encoding the given picture, so that the image-driven model can generate a driving video that is faithful to the given image.
[0050] As Figure 1 shown, the image-driven model of this embodiment mainly consists of a variational image autoencoder (VAE image encoder), a CLIP text encoder, a raw input module, several UNet modules, several temporal modules, and a variational image decoder (VAE image decoder). In the process of generating a video from text, the raw input module inputs several frames of random noise; the text prompt is encoded by the CLIP text encoder, and the obtained text encoding is injected into the UNet module of the network, so that the text encoding can interfere with the output of the model according to the semantic information of the text; among them, the UNet module processes different video frames frame by frame, and the temporal module is used to align different video frames to make the output video smooth and flicker-free; finally, through the iterative denoising process, the several frames of random noise input step by step are transformed into clear video frames that conform to the text information according to the text encoding.
[0051] In this embodiment, a conditional encoding module is added on the basis of the above image-driven model. This conditional encoding module can be inserted into the first layer of the model for generating videos from text and is compatible with the entire model for generating videos from text. Specifically, this conditional encoding module is implemented by a convolutional layer, and the structure of this convolutional layer is ; that is, this conditional encoding module uses 320 kernels of size The convolutional kernel of performs conditional encoding on the input with 4 channels, obtaining conditional encoding with the same size as the input and 320 channels; since the size and number of channels of this conditional encoding are equal to those of the output features of the original input module, the conditional encoding obtained by this conditional encoding module can be directly added to the output of the original input module, that is, this conditional encoding module is compatible with the entire text-to-video generation model. As Figure 1 shown, the input of this conditional encoding module is divided into two parts: image encoding and inter-frame consistency encoding. Among them, the image encoding is obtained by encoding the given image with a VAE image encoder, and the inter-frame consistency encoding is obtained from part (2) (not elaborated here for the time being). After the image encoding and the inter-frame consistency encoding are concatenated channel by channel and input into the conditional encoding module, a 4-channel conditional input is obtained; this conditional encoding module can retain the valid information of the given conditional frame according to the inter-frame consistency information and add it to the encoding of the noise by the original input module through encoding. Through such a design, the picture information of the original input module is significantly enhanced.
[0052] In this step, by adopting a conditional encoding module and inserting it into the image-driven model, not only can the response ability of the image-driven model to text prompts be retained, but also the high-frequency detail information of the given picture can be retained by explicitly encoding the given picture, enabling the image-driven model to generate a driving video faithful to the given image. As Figure 2 shown, where (a) is the driving image when the text prompt is "Fireworks are blooming over the castle", (b) is the driving image when the text prompt is "The castle is on fire", and (c) is the driving image when the text prompt is "Lightning strikes the castle". For the same given image of the castle, this method can efficiently drive the image according to the three different prompts of "fireworks", "on fire", and "lightning".
[0053] (2)Based on the image-driven model obtained by the above extension, construct an image-driven training dataset. Each training data sample in this training dataset contains not only the given conditional frame (i.e., the given picture), the target video frame sequence (i.e., 15 pictures temporally related to the given conditional frame), but also the inter-frame consistency encoding between each target video frame and the given conditional frame. Such training data can effectively improve the training strategy of the image-driven model, enabling it to explicitly control the intensity of the dynamic effects in the generated video.
[0054] Traditional image-driven training datasets usually only contain given conditional frames and target video frames, and use such paired datasets to train image-driven models, enabling the models to generate target video frames based on the given conditional frames. In this method, we expand the training dataset, design and calculate the inter-frame consistency between each target video frame and the given conditional frame. Such a design allows the image-driven model to input both the given conditional frame and the inter-frame consistency simultaneously, and drive the given conditional frame according to the hint of the inter-frame consistency, thereby obtaining the target video frame. By explicitly providing the inter-frame consistency to hint to the image-driven model on how to drive the image, it can greatly alleviate the ambiguity problem of image driving and further improve the controllability of the model.
[0055] Specifically, for a video sequence, we first convert it from the RGB color space to the HSV color space, and calculate the 1-norm distance between each frame in the video sequence and the conditional frame in the HSV space, which is denoted as where represents the number of the current frame, represents the number of the current video sequence. This distance measures the difference in motion amplitude between each frame in the training data and the conditional frame. The larger is, the greater the difference between the two frames and the smaller the correlation; The smaller is, the smaller the difference between the two frames and the greater the correlation. However, has a large numerical range, so we further normalize this value, which is beneficial for the model's learning. Specifically, we perform statistics on the entire dataset to obtain the maximum difference value , and use this maximum value to globally normalize the inter-frame correlation:
[0056]
[0057] After this step, we obtain a value between [0,1]. We further map the value:
[0058]
[0059] where is the final encoding used to represent the inter-frame consistency between the two frames, and are the normalized hyperparameters used to adjust the value of the inter-frame consistency encoding to fall within the required range.
[0060] Thus, the augmentation of the video training dataset is completed. Each training example contains, in addition to the conditional frame and the target video frame, the inter-frame consistency encoding between each target video frame and the conditional frame. Using such a dataset to train the image-driven model can promote the image-driven model to better encode the conditional frame, thereby improving the quality of the generated video. In addition, the user can adjust the action amplitude of the special effect by modifying the input inter-frame consistency encoding. When the value of the inter-frame consistency encoding is larger, the target video frame is more similar to the conditional frame, and the action amplitude in the special effect is smaller; when the value of the inter-frame consistency encoding is smaller, the change of the target video frame compared with the conditional frame is larger, and the action amplitude in the generated special effect is larger.
[0061] (3) Fine-tune the conditional encoding module and the temporal encoding module based on the pre-trained text-to-video model, so as to obtain an image-driven model based on text prompts that highly responds to the action-related guidance in the text prompts, is faithful to the given conditional frame, and can explicitly control the intensity of the generated special effect.
[0062] As Figure 1 shown, the conditional encoding module and the temporal module are trainable modules, and the original input module and the Unet module are non-trainable modules. First, we keep the parameters of the basic model for text-to-image unchanged, that is, Figure 1 the parameters of the original input module and the UNet module in
[0063] remain unchanged, so that the text-image knowledge pre-trained on a large number of datasets can be retained; then, we fine-tune the parts marked in red, that is, the conditional encoding module and the temporal module, on the dataset obtained in step 2. Through such a training method, the conditional encoding module can extract the required picture information from the conditional frame according to the input inter-frame consistency and add it to the encoding information output by the original input module. Since the appearance-related information has been encoded into the feature map, when we train the temporal module, the temporal module can pay more attention to the video alignment related to the action.
[0064]
[0065] Among them, represents the statistic of 5%, Represents the 95% statistic. After the training is completed, the present invention can avoid setting some extremely small values by adjusting the inter-frame consistency encoding, thereby obtaining a more stable video result. At the same time, when the user adjusts the value of the inter-frame consistency within the range of 5% to 95%, the intensity of the dynamic effects in the generated video can be effectively controlled. As Figure 3 shown, when the text prompt is "Labrador dog jumping", the schematic diagrams of the dynamic effect of (a) slight, (b) moderate, and (c) intense are given. Given the same conditional frame and the same text prompt, different degrees of dynamic effect results of slight, moderate, and intense can be obtained by adjusting the input of different values of the inter-frame consistency. After the model training is completed, as Figure 1 shown, in actual applications when using the trained image-driven model to generate content, the image that we want to use to drive the generation of the video is encoded by the VAE image encoder to obtain the image encoding, and at the same time, the inter-frame consistency encoding of each frame is manually set. After the two are concatenated channel by channel, they are input into the conditional encoding module, which can be used to guide the generation of the video. Among them, the inter-frame consistency encoding needs to be set by the user according to their own needs. For example, to generate 16 video frames, 16 numbers need to be set to represent the inter-frame consistency between each frame and the conditional frame; the value range is between [0.2, 1.0]. The larger the value (usually between [0.8, 1.0]), the higher the similarity between the generated frame and the conditional frame, and the smaller the generated action effect; the smaller the value (usually between [0.2, 0.4]), the lower the similarity between the generated frame and the conditional frame, and the more obvious the generated action effect.
[0066] It has been proven feasible through quantitative comparison, qualitative testing, and user research. The results show that this method can generate image-driven results that are more faithful to the conditional frame and text prompt than existing methods.
[0067] This method has the following advantages:
[0068] (1) The conditional encoding module is compatible with the text-to-video model, can explicitly encode and retain the information and details in the conditional frame according to the input inter-frame consistency, and can generate better video segments that restore the generated image content.
[0069] (2) By explicitly encoding the conditional frame information through the conditional encoding model and fine-tuning the conditional encoding model and the temporal model at the same time, it shows better response to the text prompt and can effectively generate relevant dynamic effects.
[0070] (3) The input of the inter-frame consistency is introduced. This design can not only allow users to avoid generating low-quality pictures with overly intense animations by avoiding setting extremely small values, but also allow users to adjust the intensity of the dynamic effects of the generated video within a reasonable value range to achieve a more controllable image-driven effect.
[0071] Example 2
[0072] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the above-mentioned content generation method based on text prompts and image driving.
[0073] Example 3
[0074] This embodiment provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the content generation method based on text prompts and image driving as described in Example 1.
[0075] As mentioned above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A content generation method based on text prompts and image driving, characterized in that, Generate a video using a pre-trained image-driven model based on a given text prompt and a given image. The training process of the image-driven model includes the following steps: Obtain samples including input text, given conditional frames, a target video frame sequence, and an inter-frame consistency encoding, where the inter-frame consistency encoding is calculated based on the given conditional frames and the target video frame sequence; Encode the given conditional frames to obtain image encodings, and based on the image encodings and the inter-frame consistency encoding, obtain conditional frame features through conditional encoding; Initialize a noise frame and obtain noise features through feature extraction; Based on the conditional frame features, the noise features, and the input text, obtain an output encoding and perform denoising to obtain a new noise frame, completing this round of iteration. Repeat this step for multiple iterations; Based on the denoised output encoding after multiple iterations, obtain output video frames, and update the parameters of the image-driven model based on the target video frame sequence and the output video frames to complete the training for the sample. Among them, conditional encoding is implemented using a conditional encoding module. The conditional encoding module is inserted into the first layer of the text-to-video generation model. The conditional encoding module includes a convolutional layer with a structure of 4×3×3×320. The conditional encoding module uses 320 convolutional kernels with a size of 4×3×3 to perform conditional encoding on an input with 4 channels, obtaining a conditional encoding with the same size as the input and 320 channels.
2. The content generation method based on text prompts and image driving according to claim 1, wherein The image-driven model includes: A conditional encoding module for obtaining conditional frame features based on the image encoding and the inter-frame consistency encoding; An original input module for obtaining noise features based on the noise frame; At least one set of Unet modules and a temporal module for obtaining an output encoding based on the conditional frame features, the noise features, and the text encoding obtained by encoding the input text.
3. A content generation method based on text prompts and image driving according to claim 2, wherein The Unet module is used to process video frames frame by frame based on the text encoding, and the temporal module is used to align video frames.
4. A method for generating content based on text prompts and image driving according to claim 2, characterized in that, The original input module and the Unet module are pre-trained and do not update parameters during the training of the image-driven model.
5. A content generation method based on text prompts and image driving according to claim 1, characterized in that, The conditional frame features are obtained by concatenating the image encoding and the inter-frame consistency encoding in the channel dimension and performing conditional encoding.
6. A method for generating content based on text prompts and image driving according to claim 1, characterized in that, The calculation process of the inter-frame consistency encoding includes: Calculate the 1-norm distance between each frame in the given conditional frames and the target video frame sequence in a preset color space, and perform global normalization processing based on the maximum value in the sample set to obtain the inter-frame consistency encoding.
7. A content generation method based on text prompts and image driving according to claim 6, characterized in that, For the 1-norm distances below a% or above b% in the sample set, replace them with the 1-norm distances at a% and b% respectively.
8. A method for generating content based on text prompts and image driving according to claim 1, characterized in that The inter-frame consistency encoding is calculated using the following formula: , , , Among them, is the 1-norm distance between the th given conditional frame and the th frame sequence frame of the target video in the HSV color space. This distance measures the difference in the motion amplitude between each frame in the training data and the conditional frame. is the maximum 1-norm distance in the sample set. is the obtained inter-frame consistency encoding. and are normalized hyperparameters. represents the 1-norm distance corresponding to the 5% statistic in the sample set. represents the 1-norm distance corresponding to the 95% statistic. is defined as.
9. An electronic device, characterized in that, Including: One or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the content generation method based on text prompts and image driving as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, including one or more programs to be executed by one or more processors of an electronic device, the one or more programs including instructions for performing the text prompt and image-driven content generation method according to any one of claims 1-8.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN116681630A
Video generation method
CN116939325A