Interactable personalized video controllable generation method and device, equipment and medium
By using a 2D image appearance updater and the Temporal-SDEdit method, the difficulties in video generation caused by inconsistencies between image and text descriptions are resolved, enabling the generation of high-quality personalized videos and ensuring consistency between video appearance and motion.
Patent Information
- Application Number
- CN202411577358.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Existing technologies struggle to generate high-quality, personalized videos when image and text descriptions are inconsistent, especially when image and text descriptions do not match, making it difficult for video diffusion models to maintain temporal consistency and texture variation.
By introducing a 2D image appearance updater and the Temporal-SDEdit method, 2D and temporal information are separated, and the text-adapted image is seamlessly integrated into a pre-trained video diffusion model. Stochastic differential equations are used to guide image generation and editing, and control the video generation process.
The text and image-to-video diffusion model capabilities have been enhanced, ensuring consistency in the appearance and motion of generated videos and improving the quality of personalized video generation.
Smart Images

Figure CN119697440B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video generation technology, and in particular to an interactive, controllable method, apparatus, device, and medium for generating personalized videos. Background Technology
[0002] With the rapid development of technology, researchers have increasingly focused on diffusion models for visual generation tasks. Significant progress has been made in text-to-image (T2I) and text-to-video (T2V) generation, but generating high-quality visual content based on text remains a challenging task. This requires not only generating visually realistic, smooth, and time-consistent motion from given text and images, but also... Early developments in video diffusion models primarily focused on single guidance, either through an image or text. Models like Animation and Zero Shot concentrate on text-guided video generation, while stable video diffusion models emphasize image-guided video generation. Personalized Image Animator (PIA) extends these approaches by integrating text and image guidance, offering greater flexibility in guided video generation. However, despite its versatility, PIA struggles to produce high-quality results when the image and text descriptions are not perfectly aligned. For example, if the text description is "a girl smiling," and the corresponding image matches this description, PIA can produce satisfactory results. However, when the text is changed to "a girl playing ball," while the image remains unchanged, PIA often fails to deliver high-quality video. This inconsistency between images and text adds an extra challenge to the diffusion process, as the model must manage both temporal consistency and texture variations simultaneously. Summary of the Invention
[0003] In order to at least partially solve one of the technical problems existing in the prior art, the present invention aims to provide an interactive and controllable method, apparatus, device and medium for generating personalized videos based on a diffusion model.
[0004] The first technical solution adopted in this invention is:
[0005] An interactive and controllable method for generating personalized videos includes the following steps:
[0006] Collect personalized images as source images;
[0007] The obtained source image is input into a 2D image appearance updater for fine-tuning to obtain the text-adapted image.
[0008] Copy multiple frames from the text-adapted image and convert the copied multiple frames into the initial video frames of text-to-video (T2V);
[0009] Based on the initial video frames, the text-adapted images and text are used as input to a pre-trained video diffusion model to generate personalized videos.
[0010] Furthermore, the step of generating personalized videos by using the text-adapted images and text as input to a pre-trained video diffusion model based on the initial video frames includes:
[0011] Set the number of frames N to adapt the image to the text. Perform N copies to generate the initial video frame for the video diffusion model.
[0012] For N-dimensional random Gaussian noise ε N Sampling is performed, and the noise ε N and Merging is performed using the scheduler of the diffusion model;
[0013] Injecting the merged, noisy video frames into the intermediate diffusion step is called the hijacking step.
[0014] For different hijacking steps, generate videos with different results and obtain the hijacking interval with better effect;
[0015] The videos generated from the hijacking zone are filtered to obtain those with satisfactory appearance and movement, which are then used as the final personalized videos.
[0016] Furthermore, the step of injecting the merged, noisy video frames into the intermediate diffusion step, referred to as the hijacking step, includes:
[0017] Injecting random Gaussian noise into different hijacking steps yields noisy video frames;
[0018] Noisy video frames are matched with a pre-trained video diffusion model and used as input frames for the intermediate diffusion process.
[0019] Based on the set number of hijacking steps, these noisy video frames are injected into the specified diffusion steps to control the video generation process.
[0020] Furthermore, the process of generating videos with different results for different hijacking steps and obtaining hijacking intervals with better results includes:
[0021] Set ranges for different hijacking steps, starting from the initial diffusion step, and gradually adjust and record the results of each generation;
[0022] Multiple video samples were generated at different hijacking steps, and the appearance and motion effects of each video sample were compared.
[0023] The hijacking range with better performance is selected based on preset visual and motion evaluation indicators;
[0024] By optimizing the algorithm to pinpoint the optimal hijacking range, the generated video is guaranteed to achieve the desired personalized effect.
[0025] Furthermore, the filtering of videos generated based on the hijacking interval to obtain videos with satisfactory appearance and motion, as the final personalized video, includes:
[0026] The video samples generated based on the optimal hijacking interval will be filtered in multiple dimensions according to appearance and motion indicators to exclude videos that do not meet the user's personalized needs.
[0027] Samples that meet the appearance and motion requirements are retained as the final output video to ensure a high degree of consistency between personalized video content and visual effects.
[0028] Further, the step of inputting the obtained source image into a two-dimensional image appearance updater for fine-tuning to obtain a text-adapted image includes:
[0029] Based on the user's personalized text description, the parameters in the 2D image appearance updater are adjusted to fit the text requirements;
[0030] The fine-tuned images are filtered and selected, and the selected images are used as input for subsequent processing to achieve personalized video content appearance adaptation.
[0031] Further, the step of copying multiple frames based on the text-adapted image and converting the copied multiple frame images into the initial video frames of the text-generated video includes:
[0032] Preprocessing is performed on multiple frames of images, including resolution adjustment and image quality enhancement;
[0033] The parameters of the initial video frame are set to provide the input basis for the generation process of the diffusion model.
[0034] The second technical solution adopted in this invention is:
[0035] An interactive, controllable personalized video generation device, comprising:
[0036] The image collection module is used to collect personalized images as source images;
[0037] The text adaptation module is used to input the obtained source image into the two-dimensional image appearance updater for fine-tuning, and obtain the text-adapted image.
[0038] The image conversion module is used to copy multiple frames of images from the text-adapted image and convert the copied multiple frames of images into the initial video frames of text-to-video (T2V).
[0039] The video generation module is used to generate personalized videos based on initial video frames, using text-adapted images and text as input to a pre-trained video diffusion model.
[0040] The third technical solution adopted in this invention is:
[0041] An electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to realize an interactive, personalized, and controllable video generation method as described above.
[0042] The fourth technical solution adopted in this invention is:
[0043] A computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement an interactive, personalized, and controllable video generation method as described above.
[0044] The fifth technical solution adopted in this invention is:
[0045] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned interactive, personalized, and controllable video generation method.
[0046] The beneficial effects of this invention are: This invention separates the two-dimensional and temporal information in text and image-based video generation through a two-dimensional image appearance updater to align the two-dimensional text and the input image, thereby enhancing the capabilities of existing text and image-to-video diffusion models; In addition, this invention proposes Temporal-SDEdit, which seamlessly integrates text-adapted images into a pre-trained video diffusion model. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following description is provided with accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1This is a schematic diagram of an interactive, personalized, and controllable video generation method based on a diffusion model in an embodiment of the present invention.
[0049] Figure 2 These are the results of video generation algorithms on the AnimateBench dataset using different text adaptation methods;
[0050] Figure 3 This is the Temporal-SDEdit stage process in this embodiment of the invention;
[0051] Figure 4 This is a flowchart illustrating the steps of an interactive, personalized, and controllable video generation method according to an embodiment of the present invention. Detailed Implementation
[0052] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0053] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0054] In the description of this invention, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.
[0055] In the description of this invention, unless otherwise explicitly defined, terms such as "set up," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.
[0056] To address the problems of existing technologies, this invention proposes a method for generating high-quality and temporally smooth videos by separating two-dimensional and temporal information. In short, the goal of this invention is to resolve the inconsistency between text and images by incorporating an additional 2D image appearance updater. This appearance updater ensures that the image is adapted to the text before being input into the video diffusion model. Furthermore, to efficiently integrate the text-adapted image into the video diffusion process, this invention also extends the Stochastic Differential Equation Guided Image Generation and Editing Method (SDEdit) to a temporal version, namely Temporal-SDEdit. This adaptation allows the video diffusion model to generate videos that are more relevant to the input.
[0057] Example 1
[0058] like Figure 4 As shown, this embodiment provides an interactive and controllable method for generating personalized videos, including the following steps:
[0059] S1. Collect personalized images as source images.
[0060] S2. Input the obtained source image into the two-dimensional image appearance updater for fine-tuning to obtain the text-adapted image.
[0061] Specifically, an auxiliary, interactive 2D diffusion model is used to synchronize 2D appearance details from text and image inputs to obtain images with varying degrees of text adaptation.
[0062] As an optional implementation, step S2 specifically includes the following steps:
[0063] S21. Adjust the parameters in the 2D image appearance updater to match the text requirements based on the user's personalized text description.
[0064] S22. Filter and select the fine-tuned image, and use the selected image as input for subsequent processing to achieve personalized video content appearance adaptation.
[0065] S3. Copy multiple frames of images from the text-adapted image and convert the copied multiple frames of images into the initial video frames of text-to-video (T2V).
[0066] In some embodiments, step S3 further includes a preprocessing step for multiple frames of images, wherein the preprocessing includes resolution adjustment, image quality enhancement, etc. Additionally, by setting the parameters of the initial video frames, an input basis is provided for the generation process of the diffusion model.
[0067] S4. Based on the initial video frames, the text-adapted image and text are used as input to a pre-trained video diffusion model to generate a personalized video.
[0068] As an optional implementation, step S4 specifically includes the following steps:
[0069] S41. Set the frame number N and adapt the image to the text. Perform N copies to generate the initial video frame for the video diffusion model.
[0070] S42, Regarding N-dimensional random Gaussian noise ε N Sampling is performed, and the noise ε N and Merging is performed using a scheduler based on a diffusion model.
[0071] S43. Injecting the merged, noisy video frames into the intermediate diffusion step is called the hijacking step.
[0072] In some embodiments, step S43 includes the following steps:
[0073] S431: Inject random Gaussian noise into different hijacking steps to obtain noisy video frames;
[0074] S432: Match noisy video frames with a pre-trained video diffusion model as input frames for the intermediate diffusion process;
[0075] S433: Based on the set number of hijacking steps, inject these noisy video frames into the specified diffusion steps to control the video generation process.
[0076] S44. Generate videos with different results for different hijacking steps, and obtain the hijacking interval with better results.
[0077] In some embodiments, step S44 includes the following steps:
[0078] S441: Set a range for different hijacking steps, starting from the initial diffusion step, gradually adjust and record the results generated each time;
[0079] S442: Generate multiple video samples under different hijacking steps, and compare the appearance and motion effects of each video sample;
[0080] S443: Select the hijacking interval with better performance based on preset visual and motion evaluation indicators;
[0081] S444: By optimizing the algorithm to lock in the optimal hijacking interval, the generated video is ensured to achieve the desired personalized effect.
[0082] S45. Filter the videos generated based on the hijacking interval to obtain videos with satisfactory appearance and motion, which will be used as the final personalized videos.
[0083] In some embodiments, step S45 includes the following steps:
[0084] S451: Based on the video samples generated from the optimal hijacking interval, the generated video samples are screened in multiple dimensions according to appearance and motion indicators to exclude videos that do not meet the user's personalized needs.
[0085] S452: Retain samples that meet the appearance and motion requirements as the final output video to ensure a high degree of consistency between the personalized video content and the visual effects.
[0086] In this embodiment, Temporal-SDEdit is proposed to seamlessly integrate the refined image into a pre-trained video diffusion model, thereby better maintaining the consistency of T2V video generation and using a denoising process with a smaller time step than the original. Finally, a video with satisfactory appearance and motion is generated.
[0087] The following detailed explanation is provided in conjunction with the accompanying drawings and specific embodiments.
[0088] This embodiment provides a controllable generation method for 2D interactive personalized videos based on a diffusion model. It theoretically verifies that controllable guidance from different sources brings additional complexity, often increasing the burden on the model and making it difficult to maintain semantic and temporal consistency. Essentially, it enhances the capabilities of existing diffusion models for text, images, and videos.
[0089] To address the misalignment issue between input text and images in video generation, a text and an image are required as input conditions. To alleviate the burden on the model, we separate the 2D and temporal information of the text and image-based video generation using a 2D image appearance updater. This process adjusts the image in a given text-image pair to enhance coherence. Subsequently, we introduce Temporal-SDEdit to seamlessly integrate the refined input image into the video diffusion model.
[0090] In this embodiment, the overall process is as follows: Figure 1 As shown, firstly, the collected personalized images, used as source images, are defined by the text-based T2I diffusion model and then adapted using the same T2I model. In the first step, the text-adapted images alleviate the burden of appearance updates throughout the complex video generation process. Secondly, in the second step, we copy the 'F' frame number from the text-adapted images and convert it into the initial video frames for T2V. Furthermore, we use "Temporal-SDEdit" to better maintain the consistency of T2V video generation, employing a denoising process with a smaller time step than the original. Finally, a video with satisfactory appearance and motion is generated.
[0091] The details of the method proposed in this embodiment will be discussed from two aspects, mainly including the 2D image appearance update stage process and the Temporal-SDEdit stage process.
[0092] (1) 2D image appearance update process
[0093] The goal of this process is to update the input image to a text-adapted image. This can be achieved by using an additional 2D image appearance updater. Popular 2D image appearance updaters include Prompt-to-Prompt and InstructPix2Pix. InstantID uses only a single face image to personalize images in various styles. DragDiffusion is an interactive point-based image update. In this embodiment, a two-dimensional diffusion model is primarily used to update the image to obtain the desired text-adapted image. We choose the commonly used InstructPix2Pix as our 2D image appearance updater, which offers rich instruction-level prompts, including objects, faces, backgrounds, etc. The process specifically includes the following steps:
[0094] Step 1: Personalized image collection for users.
[0095] Step 2: Obtain the text-adapted image. First, input the collected images into a 2D appearance updater (diffusion model) for fine-tuning to obtain the text-adapted image, such as... Figure 2 As shown.
[0096] Step 3: Select and filter randomly generated adapted images from multiple sources. Considering the uncertainty brought about by randomness, establish a one-way mapping relationship between the prompt of the source image and the prompt of the adapted text, and finally obtain the most satisfactory image.
[0097] Step 4: Video Generation. Using both the text-adapted image and text as input, a video diffusion model is used for fine-tuning to generate a personalized video.
[0098] (2) Temporal-SDEdit Stage Flow
[0099] The core idea of SDEdit is to "hijack" normal noise into the image generation process by injecting noisy images into intermediate diffusion steps. Formally, instead of gradually denoising the random noise ε to the final diffusion step and generating χ0, we use... To replace ε, where It is a real image (In our example, text adapts to the image) and a combination of noise ε. By "hijacking" the diffusion process, the SDEdit method guarantees that the unconditional diffusion model applies to the real image. Image generation conditions. In this embodiment, we extend the SDEdit method to temporal-SDEdit. See [link to documentation]. Figure 3 The process specifically includes the following steps:
[0100] Step 1: Given N frames, we repeat the text-adapted image. Perform N times to generate
[0101] Step 2: We apply N-dimensional random Gaussian noise ε N Sampling is performed, and the noise ε N and Merge with the diffusion scheduler.
[0102] The third step: Injecting the merged noise into the intermediate diffusion step, is called the hijacking step.
[0103] Step 4: Different hijacking steps will produce different video results, and finally, a hijacking range with better results can be obtained.
[0104] Step 5: Filter the videos generated from the hijacking area to obtain videos with satisfactory appearance and motion.
[0105] In summary, compared with the prior art, the present invention has at least the following advantages and beneficial effects:
[0106] (1) Fast robust perception capability: The system framework of this invention has a good ability to understand text and image alignment of diffusion models.
[0107] (2) Quick deployment and fine-tuning: This method only requires collecting images and fine-tuning, and is highly reproducible, which can lay the foundation for subsequent real-time testing.
[0108] (3) Flexibility and scalability: This framework can be used in any scenario. It can be customized to adapt to different scenario requirements by first deploying it in a fine-tuning manner.
[0109] Example 2
[0110] This embodiment provides an interactive, controllable personalized video generation device, including:
[0111] The image collection module is used to collect personalized images as source images;
[0112] The text adaptation module is used to input the obtained source image into the two-dimensional image appearance updater for fine-tuning, and obtain the text-adapted image.
[0113] The image conversion module is used to copy multiple frames of images from the text-adapted image and convert the copied multiple frames of images into the initial video frames of text-to-video (T2V).
[0114] The video generation module is used to generate personalized videos based on initial video frames, using text-adapted images and text as input to a pre-trained video diffusion model.
[0115] Since this device is an interactive, personalized, and controllable video generation device according to an embodiment of the present invention, and the principle by which this device solves the problem is similar to that of this method, the implementation of this device can refer to the implementation process of the above-described method embodiment, and repeated details will not be repeated.
[0116] Example 3
[0117] This invention also provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set, or an instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to achieve the following: Figure 4 This illustrates an interactive and controllable method for generating personalized videos.
[0118] It is understood that the memory may include random access memory (RAM) or read-only memory. Optionally, the memory may include non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a stored program area and a stored data area, wherein the stored program area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the stored data area may store data created according to the use of the server, etc.
[0119] A processor may include one or more processing cores. The processor connects to various parts of the server via various interfaces and lines, executing instructions, programs, code sets, or instruction sets stored in memory, and accessing data stored in memory to perform various server functions and process data. Optionally, the processor may be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor may integrate one or more of the following: Central Processing Unit (CPU) and Modem. The CPU primarily handles the operating system and applications; the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor.
[0120] Since this electronic device is an electronic device corresponding to the interactive and controllable generation method for personalized videos in this embodiment of the invention, and the principle of solving the problem by this electronic device is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0121] Example 4
[0122] This invention also provides a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to achieve the following: Figure 4 This illustrates an interactive and controllable method for generating personalized videos.
[0123] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0124] Since this storage medium is the storage medium corresponding to an interactive, personalized, and controllable video generation method according to an embodiment of the present invention, and the principle of solving the problem by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and repeated parts will not be described again.
[0125] Example 5
[0126] In some possible implementations, various aspects of the methods of the embodiments of the present invention can also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps of an interactive, personalized, controllable video generation method according to various exemplary embodiments of this application as described above. The executable computer program code or "code" for performing the various embodiments can be written in high-level programming languages such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0127] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0128] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0129] The above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made based on the essence of the content of the present invention should be covered within the scope of protection of the present invention.
Claims
1. An interactive personalized video controllable generation method, characterized in that, The method comprises the following steps: Collecting personalized images as source images; Inputting the obtained source images into a two-dimensional image appearance updater for fine-tuning to obtain text-adapted images; Copying multiple frames of images from the text-adapted images, and converting the copied multiple frames of images into initial video frames of the text-to-video video; Based on the initial video frames, inputting the text-adapted images and the text into a pre-trained video diffusion model to generate a personalized video through a Temporal-SDEdit process, comprising: Set frame number Image adapted to text Carrying out Secondary copying to produce initial video frames of the video diffusion model ; To with random Gaussian noise sampled and the noise and merged with a scheduler using a diffusion model; Injecting the merged video frames with noise into intermediate diffusion steps, referred to as hijacking steps; For different hijacking steps, generate videos with different results and obtain a hijacking interval with better effects; Filtering the videos generated according to the hijacking interval to obtain a video with satisfactory appearance and motion as the final personalized video; The step of generating videos with different results for different hijacking steps and obtaining a hijacking interval with better effects comprises: Setting a range for different hijacking steps, starting from the initial diffusion step, and gradually adjusting and recording the results generated each time; Generating multiple video samples under different hijacking steps, and comparing the appearance and motion effects of each video sample; According to the preset visual and motion evaluation index, the hijacking interval with better effect is screened out; Locking the best hijacking interval through an optimization algorithm to ensure that the generated video achieves the expected personalized effect; The step of filtering the videos generated according to the hijacking interval to obtain a video with satisfactory appearance and motion as the final personalized video comprises: According to the appearance and motion indicators, the video samples generated based on the best hijacking interval are subjected to multi-dimensional screening, and the videos that do not meet the user's personalized needs are excluded; The samples that meet the appearance and motion requirements are retained as the final output video to ensure the high consistency of the video content personalization and visual effect.
2. The method of claim 1, wherein, The step of injecting the merged video frames with noise into intermediate diffusion steps, referred to as hijacking steps, comprises: Injecting random Gaussian noise for different hijacking steps to obtain noisy video frames; Matching the noisy video frames with the pre-trained video diffusion model as the input frames of the intermediate diffusion process; According to the set hijacking steps, the noisy video frames are injected at the specified diffusion steps to control the video generation process.
3. The method of claim 1, wherein, The step of inputting the obtained source images into a two-dimensional image appearance updater for fine-tuning to obtain text-adapted images comprises: Adjusting the parameters in the two-dimensional image appearance updater to adapt to the text requirements according to the user's personalized text description; Filtering and selecting the fine-tuned images, and taking the selected images as the input of the subsequent processing process to realize the appearance adaptation of the personalized video content.
4. The method of claim 1, wherein, The step of copying multiple frames of images from the text-adapted images, and converting the copied multiple frames of images into initial video frames of the text-to-video video comprises: Pretreating the multiple frames of images, wherein the pretreatment includes resolution adjustment and image quality enhancement; Setting parameters for the initial video frames to provide input basis for the generation process of the diffusion model.
5. An interactive personalized video controllable generation apparatus for implementing the method of any one of claims 1 to 4, characterized in that, The method comprises the following steps: An image collection module is configured to collect personalized images as source images; The text adaptation module is configured to input the obtained source image into a two-dimensional image appearance updater for fine-tuning to obtain a text-adapted image. The image conversion module is configured to copy multiple frames of images from the text-adapted image, and convert the copied multiple frames of images into initial video frames of a text video. The video generation module is configured to input the text-adapted image and the text into a pre-trained video diffusion model based on the initial video frames, and generate a personalized video.
6. An electronic device, comprising: The electronic device includes a processor and a memory, and the memory stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the method of any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the method of any one of claims 1-4.
Citation Information
Patent Citations
Video generation method and device, model training method and device, equipment and medium
CN116320216A
Image extension method and device based on generative image model, equipment and medium
CN117745522A
Image processing method, device, equipment, medium and program product
CN118711113A