Zero sample text driven video editing method based on diffusion model

Through the zero-sample text-driven video editing method based on the diffusion model, the problems of video timing maintenance, content fidelity and calculation complexity in the prior art are solved, and efficient and convenient video editing effects are achieved.

CN120091182APending Publication Date: 2025-06-03BEIJING INFORMATION SCI & TECH UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510232326.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing video editing technologies have challenges in handling video timing maintenance, content fidelity and computational complexity, especially in the complex temporal information modeling within videos and the need for large amounts of labeled video data.

Method used

The zero-sample text-driven video editing method based on the diffusion model is adopted to realize video editing through video frame transformation to potential space, semantic fusion of inter-frame diffusion features, self-attention guidance and foreground partial diffusion.

Benefits of technology

This method can maintain high fidelity and consistency while reducing calculation time and calculation amount, accurately present the expected editing effect, and improve the efficiency and performance of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091182A_ABST
    Figure CN120091182A_ABST
Patent Text Reader

Abstract

The invention provides a zero-sample text-driven video editing method based on a diffusion model, which belongs to the technical field of video editing and comprises the following steps: converting a video frame into a potential space; semantic fusion of inter-frame diffusion features is carried out; performing self-attention guidance; and diffusing the foreground part. According to the invention, video frames are converted to a potential space; semantic fusion of inter-frame diffusion features is carried out; by means of self-attention guiding and foreground part diffusion, high fidelity and consistency can be kept, the expected editing effect can be accurately presented in the examples, and the calculation time and the calculation amount are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video editing, and in particular, to a zero-sample text-driven video editing method based on a diffusion model. Background Art

[0002] Video, as a major media form, plays an important role in social life and has a very wide range of applications. In a large number of video applications, it is often necessary to edit and adjust videos. For example, adjusting the clothing of a person in the video to another style or color, adjusting the scene of the video background, replacing a certain target in the video with another target, and so on. Traditional video editing mainly relies on video editing software through manual methods. This video editing method is often time-consuming and laborious, and requires a large amount of cost investment. With the rapid development of artificial intelligence generative models recently, using artificial intelligence for automatic video editing can greatly improve the efficiency and performance of video editing. This application is under this demand, attempting to edit videos through language description to achieve a more convenient, efficient, and realistic video editing effect.

[0003] Generative models come in various forms, among which diffusion models have relatively the most detailed and realistic effects, thus gaining extensive attention. Among the currently published diffusion models, the models represented by Stable Diffusion are more prominent. These models demonstrate excellent capabilities in generating high-quality image content and enabling multimodal conversion. To further enhance the attractiveness and practical application value of diffusion models, the present invention extends them to the field of video editing, achieving desired video editing effects through the correction and adjustment of specific regions of the video. Although impressive results have been achieved in image generation based on diffusion models, video generation still faces many challenges. These challenges mainly include the complex temporal information modeling within the video and the need for a large amount of annotated video data. These problems bring new difficulties and opportunities for further research in this field. Video editing is regarded as a process of using video generation techniques to combine and optimize video, audio, and image materials to create a final video work that is coherent, attractive, and meets the expected effects. Especially after the introduction of deep learning models, significant progress has been made in related technologies. In text-driven video editing tasks, the user provides an input video I = {I^i|i ∈ [1, n]}, along with the target text description (source prompt c). The core objectives of this task include three points: 1. Video semantic consistency, that is, the generated video should strictly correspond to the semantics of the target prompt c; 2. Fidelity of video content, that is, the content of each frame of the edited video should be as consistent as possible with the corresponding frame of the original video; 3. Temporal coherence, that is, the generated video should show consistency and smoothness on the time axis. Although it seems easy to achieve video editing through frame-by-frame image editing techniques, even using content-preserving diffusion inversion methods (such as DDIM inversion) and cross-attention guidance, serious flickering will occur between the generated video frames, such as Figure 1As shown in the second line of , due to the lack of temporal modeling, frame-by-frame image editing produces temporally inconsistent results when alternating video styles. Previous studies have partially addressed these issues by expanding 2D videos into pseudo-3D representations and extending spatial attention to spatio-temporal self-attention. However, these strategies are still unsatisfactory in terms of video temporal preservation and content fidelity. In addition, when the editing target involves small objects in the video, existing methods often have side effects on changing other irrelevant factors. For example, when editing small target objects, it may unnecessarily modify the features of the overall scene, such as the background, lighting, and texture, so that the modification of the color and shape of small objects will affect the background color change or the texture of other objects, which significantly limits the practical application effect and feasibility of video editing technology. In addition, special attention must be paid to maintaining temporal consistency between frames when editing videos. If the inter-frame correlation is not fully considered, the edited video may appear incoherent or even unrealistic. To address this challenge, existing research has introduced methods such as diffusion inversion and cross-frame self-attention to enhance temporal consistency. However, these methods are usually accompanied by a significant increase in computational burden and still need to be further optimized for more efficient applications.

[0004] Therefore, there is an urgent need in the art for a technical solution that can reduce computational complexity and improve editing efficiency.

[0005] The information disclosed in this background section is only intended to increase the understanding of the overall background of the present invention and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention

[0006] The object of the present invention is to provide a video editing method that can reduce computational complexity and improve editing efficiency.

[0007] To achieve the above object, the present invention provides the following solution:

[0008] A zero-sample text-driven video editing method based on a diffusion model, comprising:

[0009] Transforming video frames into the latent space;

[0010] Semantic fusion of inter-frame diffusion features;

[0011] Self-attention guidance;

[0012] Foreground part diffusion.

[0013] Optionally, the transforming video frames into the latent space includes:

[0014] Using DDIM inversion to map video frames from the pixel space to the latent representation in the noise space to generate an inversion trajectory

[0015] By optimizing the null-text embedding trajectory Learn soft text embeddings that are more closely aligned with the video content; the optimization goal is to minimize the difference between the null-text embedding and the DDIM inversion trajectory, i.e.:

[0016]

[0017] where f θ represents the operation of applying the DDIM sampling process on the basis of the pre-trained diffusion model ∈ θ Based on the above, apply the DDIM sampling process

[0018] Optionally, the semantic fusion of the inter-frame diffusion features includes:

[0019] Apply the DDIM reverse process to each frame of the input video to obtain a series of latent variables During the editing process, extract the query vector Q θ of each frame from the cross-attention module of each layer of the model network ∈ i and the key-value vectors K i , V i ; Concatenate the query vectors Q i of all frames, and obtain an attention map that fuses the information of all frames by calculating the attention mechanism

[0020] Optionally, the self-attention guidance includes:

[0021] During the reconstruction process, sample using the inversion result of the original video and the source prompt, and at the same time inject the self-attention maps generated in the Unet attention module, and these self-attention maps are calculated by the Softmax function

[0022] Optionally, the foreground partial diffusion includes:

[0023] First, perform DDIM inversion and noise addition on the video frame latent vector z 0 , and at the same time execute the Null-text optimization algorithm to obtain the null-text vector and stop at time step T t ; Then, extract the foreground latent vector according to the foreground mask M and continue to perform DDIM inversion and noise addition on it to obtain a set of frame noise maps A set of null-text vectors and a set of foreground region noise maps During the reconstruction process, perform reverse denoising, use the null-text vector for optimization, and extract the self-attention map M t , during editing, for the foreground region noise map Perform sampling at time T t Stop at the time step and combine with the frame noise map Fuse them, then perform the sampling operation, and use the empty text vector for optimization and inject the self-attention map M t , and finally obtain the edited video.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] The present invention provides a zero-sample text-driven video editing method based on a diffusion model. By transforming video frames into the latent space, semantic fusion of inter-frame diffusion features, self-attention guidance, and foreground part diffusion, it can not only maintain high fidelity and consistency, but also accurately present the expected editing effects in these instances, and reduce the computing time and amount of computation. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a schematic diagram showing the results of existing frame-by-frame video editing.

[0028] Figure 2 It is a flowchart of the method provided by the embodiment of the present invention.

[0029] Figure 3 It is a comparison diagram with mainstream models provided by the embodiment of the present invention.

[0030] Figure 4 It is a ablation comparison diagram of inter-frame diffusion feature fusion provided by the embodiment of the present invention.

[0031] Figure 5 It is a self-attention guidance ablation comparison diagram provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0033] The object of the present invention is to provide a video editing method capable of reducing computational complexity and improving editing efficiency.

[0034] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Embodiment 1:

[0036] This embodiment provides a zero-sample text-driven video editing method based on a diffusion model, as Figure 2-3 shown, including:

[0037] Transforming video frames into the latent space;

[0038] Semantic fusion of inter-frame diffusion features;

[0039] Self-attention guidance;

[0040] Foreground part diffusion.

[0041] In one embodiment, DDIM inversion is used to map video frames from the pixel space to the latent representation in the noise space to generate an inversion trajectory However, the latent noise obtained by DDIM inversion may not be completely aligned with the text description provided by the user, resulting in a difference between the reconstructed video and the original video. To this end, a prompt tuning method is introduced to learn soft text embeddings that are more closely aligned with the video content by optimizing the empty text embedding trajectory The optimization objective is to minimize the difference between the empty text embedding and the DDIM inversion trajectory, that is:

[0042]

[0043] where f θ represents the operation of applying the DDIM sampling process on the basis of the pre-trained diffusion model ∈ θ .

[0044] In video editing, directly using image editing methods for each frame independently may lead to inconsistencies in the content between frames. To solve this problem, this embodiment proposes a semantic fusion method for inter-frame diffusion features. First, apply the DDIM reverse process to each frame of the input video to obtain a series of latent variables During the editing process, extract the query vector Q θ and the key-value vectors K i and V i of each frame from the cross-attention module of each layer of the model network ∈ i ; concatenate the query vectors Q i of all frames, and obtain an attention map that fuses the information of all frames by calculating the attention mechanism This operation enables the query information of each key frame to be shared with all other frames, thus maintaining the appearance consistency of the edited frames. The final output φ(I i ) is obtained by multiplying the attention map with the value vector V i . This not only preserves the hierarchical structure of the original video features but also ensures the consistency between frames.

[0045] In video editing tasks, to maintain the similarity between the edited video and the original video, a self-attention guidance method is introduced. This method guides the editing process by leveraging the self-attention maps in the model, which contain the spatial structure information of the original video, to ensure that the editing results retain as much detail and structure of the original video as possible. Specifically, in the reconstruction process, the inversion result of the original video and the source prompt are used for sampling, while injecting the self-attention maps generated in the Unet attention module. These self-attention maps are calculated through the Softmax function.

[0046] Since a video consists of multiple frames and each frame needs to be sampled, the computational cost is huge, and the temporal consistency between frames also needs to be maintained, further increasing the computational complexity. To address this issue, this paper proposes a foreground region partial diffusion method to preferentially process the foreground region to reduce the computational overhead and time. The specific steps are as follows:

[0047] First, perform DDIM inversion and noise addition on the video frame latent vector z 0 , and at the same time execute the Null-text optimization algorithm to obtain the null text vector and stop at the T t time step; then, extract the foreground latent vector according to the foreground mask M and continue to perform DDIM inversion and noise addition on it to obtain the frame noise map set null text vector set and the foreground region noise map set During the reconstruction process, perform reverse denoising, use the null text vector for optimization, and extract the self-attention map M t . When editing, sample the foreground region noise map and stop at the T t time step. Combine with the frame noise map , then perform the sampling operation. During this period, use the null text vector for optimization and inject the self-attention map M t , and finally obtain the edited video.

[0048] As Figure 4As shown, in this embodiment, the method is qualitatively compared with a variety of other methods. It can be seen that after 300 epochs of training, the objects generated by Tune-A-Video are slightly different from the original objects, indicating that this method is less effective in maintaining the texture structure of the original image objects. For the video generated by Render-A-Video, although the structural information of the objects is retained, the texture does not match well with the text, and there are significant differences between the surrounding background and the overall style and the original video. However, the method of this embodiment demonstrates excellent performance, not only being able to maintain high fidelity and consistency, but also accurately presenting the expected editing effects in these instances, and reducing the computing time and amount of computation.

[0049] Figure 4 The following shows the effect comparison of applying the self-attention guidance method in the video editing method proposed in the embodiment of the present invention. The first row above shows the original video frame, which is used to provide a reference benchmark; the second row shows the editing result generated without self-attention guidance; the third row shows the generated result after adding self-attention guidance. From the comparison, it can be seen that the result without self-attention guidance shows a certain degree of deviation or blur in the target area (such as within the red border), and the error between adjacent frames is relatively obvious, with more jitter from the perspective of the video; while the foreground objects between adjacent frames generated by the method with self-attention guidance are more consistent in structural contours, with smaller errors, and are smoother from the perspective of the video.

[0050] In addition, Figure 5 The following shows the effect comparison of applying the self-attention guidance method in the video editing method proposed in this embodiment: Similarly, the first row above shows the original video frame, which is used to provide a reference benchmark. In the case of no self-attention guidance (the second row), the generated fox has significant differences in shape and contour from the fox contour in the original video frame, with obvious deviations in the ears, facial details, and body contour, and overall lacks a faithful restoration of the details of the original video frame. While the effect with self-attention guidance (the third row) shows that the fox contour is closer to the fox in the original video frame, the ears, facial expressions, and body shape are more in line with the characteristics of the original picture, and the dynamic consistency is significantly improved, and the shape and posture are more stable between consecutive frames. Comparative analysis shows that self-attention guidance plays a key role in contour preservation and detail restoration, significantly improving the matching degree between the generated content and the original video frame, effectively reducing shape deviation, and is of great significance for enhancing the authenticity and consistency of the generated content in video editing.

[0051] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0052] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A zero-shot text-driven video editing method based on a diffusion model, characterized in that: include: Transform video frames into latent space; Semantic fusion of inter-frame diffusion features; Self-attention guidance; The foreground is partially diffused.

2. The zero-shot text-driven video editing method based on diffusion model according to claim 1, characterized in that: The transformation of the video frame to the latent space includes: DDIM inversion is used to map the video frame from pixel space to the latent representation in voice space, generating an inversion trajectory. Embedding tracks by optimizing empty text Learn soft text embeddings that are more closely aligned with the video content; the optimization goal is to minimize the difference between the null text embedding and the DDIM inversion trajectory, i.e.: Among them, f θ Represents the pre-trained diffusion model ∈ θ Based on the operation of applying DDIM sampling process.

3. The zero-shot text-driven video editing method based on diffusion model according to claim 1, characterized in that: The semantic fusion of the inter-frame diffusion features includes: Apply the DDIM inverse process to each frame of the input video to obtain a series of latent variables During the editing process, from the model network ∈ θ The query vector Q of each frame is extracted in each layer of the cross attention module i and the key-value vector K i , V i ; The query vector Q of all frames i Splice them together and get an attention map that integrates all frame information by calculating the attention mechanism 4. The zero-shot text-driven video editing method based on diffusion model according to claim 1, characterized in that: The self-attention guidance includes: During the reconstruction process, the inversion results and source cues of the original video are used for sampling, and the self-attention maps generated in the Unet attention module are injected, which are calculated by the Softmax function.

5. The zero-shot text-driven video editing method based on diffusion model according to claim 1, characterized in that: The foreground portion of the diffusion includes: First, perform DDIM inversion and noise addition on the video frame potential vector z0, and execute the Null-text optimization algorithm to obtain the empty text vector And in T t The time step stops; then, the foreground potential vector is extracted according to the foreground mask M And continue to perform DDIM inversion and noise addition to obtain a frame noise map set Empty text vector set And the foreground area noise map collection During the reconstruction process, reverse denoising is performed, using an empty text vector Optimize and extract from the attention map M t , when editing, the noise map of the foreground area Perform sampling at T t The time step stops and Frame Noise Fusion, then sampling operation, using empty text vector Optimize and inject self-attention map M t , and finally get the edited video.

Citation Information

Patent Citations

  • Video generation model training method and device, equipment and storage medium

    CN117499711A

  • Zero sample video editing method based on diffusion model

    CN118037569A

  • Video coloring method and system based on diffusion model, and storage medium

    CN118101912A

  • Video editing method based on diffusion model corresponding relation

    CN118695042A

  • Null-text inversion for editing real images using guided diffusion models

    WO2024107884A1