Controllable video editing method based on structure-appearance information fusion

By integrating structure and appearance information in video editing, the inaccuracy problem caused by video editing relying on text control in the prior art is solved, and fine control and consistency of the appearance of video is achieved, and high-quality and diverse videos are generated.

CN120017907APending Publication Date: 2025-05-16FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510083547.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing video editing methods rely on text control, resulting in inaccurate results, requiring manual correction, and it is difficult to maintain the consistency and structural integrity of video motion.

Method used

The controllable video editing method based on structure-appearance information fusion is adopted. The structure information of the video is extracted through the structural condition control network, and the appearance condition control network is combined with the appearance condition control network to extract appearance details from the edited image. After the fusion, the edited video is generated through the video editing backbone network.

Benefits of technology

It realizes fine control of the appearance of video, enhances the accuracy and consistency of video editing, and can generate videos with a wide range of structures, styles and appearances to meet different visual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017907A_ABST
    Figure CN120017907A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, and particularly discloses a controllable video editing method based on structure-appearance information fusion. The method comprises the following steps: introducing ControlNet as a structure condition control network, extracting and injecting various structure information from an input video, then introducing an appearance condition control network for combining an image edited by a user as appearance control information in a video editing process, constructing a video editing main frame based on AnimateDiff, and carrying out video editing on the basis of the AnimateDiff. And fusing the multi-scale structure information feature map and the multi-scale appearance information feature map, and generating an edited video in combination with input text information. Compared with the prior art, a flexible editing tool is provided by coordinating the appearance information and the structure information. A user can edit a video according to specific requirements in combination with various pre-trained personalized text-image generation models to generate videos of various styles, structures and appearances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a controllable video editing method based on structure-appearance information fusion. Background Art

[0002] In recent years, the field of visual content generation has undergone profound changes, mainly due to the rapid development and widespread application of diffusion generation models. Based on these innovations, the Text-to-Image (T2I) model came into being and became a revolutionary paradigm, allowing users to generate images only through text descriptions. Among them, the famous T2I model, Stable Diffusion, has become the cornerstone of the field of image generation with its outstanding capabilities. In addition, the introduction of ControlNet further expands the capabilities of image editing. ControlNet introduces diverse structured guidance in the image generation process, allowing the Stable Diffusion model to more accurately control the image generation process. This not only improves the accuracy of editing, but also enables users to better control the overall structure and details of the image during the editing process, thereby achieving more complex and high-quality image editing effects.

[0003] Previous editing methods mainly focus on the field of image editing, aiming to achieve different visual effects by modifying existing images. The core goal of these methods is to adjust and improve images based on the input text description and the attached conditions, so that the generated images are more in line with the needs and expectations of users. Despite the remarkable progress in the field of image editing, video editing still faces great challenges. First, image generation models often ignore the continuity of temporal information, such as maintaining the consistency of video motion. Directly applying image editing methods to videos may lead to obvious flickering defects. Secondly, the lack of large-scale text-video datasets also brings difficulties to the field of video editing. It is very challenging to develop a general video editing model similar to Stable Diffusion in the field of image generation. Finally, during the inversion process of the video, due to error accumulation, the inverted noise may destroy the motion and structure of the original video.

[0004] Unlike image editing, video editing requires not only adjusting the appearance of frames, but also ensuring the consistency between frames to maintain the quality of the video, which makes video editing a more challenging task. Current video editing methods are generally divided into two categories: inversion-based methods and inversion-free methods. 1. Inversion-based methods: These methods use DDIM inversion to transform the original video into a latent variable space, and then generate the edited video through a denoising process. Usually, these methods use attention features during the inversion process to ensure structural consistency with the original video. For example, DDIM inversion is performed on the original video, and the fine-tuned parameters trained on the video are used for editing; the correspondence between frame features calculated during the inversion process is used. 2. Inversion-free methods: This type of method avoids the use of DDIM inversion and mainly relies on ControlNet to preserve the structural information of the original video. For example, ControlNet is integrated into the video generation process, and full cross-frame attention and interlaced frame smoothing techniques are used. However, these methods mainly rely on input text guidance for video editing, often lacking direct control over the appearance of the generated video, resulting in inaccurate results, and users are therefore required to perform complex manual modifications to the text prompts to meet their preferences. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art that only relies on text to control the appearance of the video, resulting in inaccurate results requiring manual correction, and to bridge the gap between video editing and image editing, and to provide a controllable video editing method based on structure-appearance information fusion.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] A controllable video editing method based on structure-appearance information fusion, the method controls and edits the video to be edited based on a text prompt input by a user, and the specific steps include:

[0008] Extract structural information from the video to be edited and input it into the structural condition control network based on ControlNet to obtain a multi-scale structural information feature map;

[0009] Edit the first frame of the video to be edited and input it into the appearance condition control network based on SparseCtrl to obtain a multi-scale appearance information feature map;

[0010] The multi-scale structural information feature map and the multi-scale appearance information feature map are fused and passed into the video editing backbone network based on AnimateDiff. The video editing backbone network generates an edited video based on the input text information.

[0011] As a preferred technical solution, the structural condition control network uses ControlNet to encode different structural information frame by frame, and the encoded structural information is integrated into the decoder of the video editing backbone network to control the video generation process. The entire process of the structural condition control network is expressed as follows:

[0012] y s =F1(x)+α s *Z1(F2(x+Z2(c s )))

[0013] Among them, F1 represents the encoder block of the backbone network; F2 represents the encoder part of the structural conditional control network ControNet corresponding to the backbone network encoder F1, Z1 and Z2 represent two zero-initialized convolutional layers; α s is a hyperparameter that controls the structural conditions and network strength; x represents the shape The video in the latent variable space of c s represents the structure representation extracted from the input video; y s Represents the output feature map containing structural information.

[0014] As a preferred technical solution, the structural information includes depth, posture, line draft and edge lines.

[0015] As a preferred technical solution, the structural condition control network constructs a ControlNet corresponding to each type of structural information; whenever a type of structural information is input, the corresponding ControlNet is used to extract a multi-scale structural information feature map.

[0016] As a preferred technical solution, the method uses an image editing tool to edit the first frame of the video to be edited to obtain an appearance control picture; a frame of all 0s is spliced ​​after the appearance control picture, and is connected with a conditional mask of all 1s in the first frame and all 0s in subsequent frames in the channel dimension and then input into an appearance conditional control network.

[0017] As a preferred technical solution, the appearance condition control network includes an appearance encoder and an appearance information propagation module;

[0018] The appearance encoder expands each convolution layer and attention layer from 2D to pseudo 3D layer based on ControlNet, and extracts appearance condition information from the appearance control picture;

[0019] An appearance propagation module is incorporated after each appearance encoder, and the appearance information propagation module uses a temporal attention layer to propagate the appearance condition information extracted by the appearance encoder to each frame of the video.

[0020] As a preferred technical solution, the processing process of the appearance condition control network is as follows:

[0021] y a =F3(x)+α a *Z3(F4(Z4(Concat(I,M))))

[0022] Among them, y a represents a feature map with appearance information; F3 is an encoder module of the video editing backbone network, F4 is an encoder module of the appearance condition control network, Z3 and Z4 are two zero-initialized convolutional layers; α a is a hyperparameter controlling the strength of the appearance conditional control network1; x represents the video in the latent variable space; I is a multi-frame input containing edited frames and several zero maps; M contains a binary conditional mask indicating the edited image frame.

[0023] As a preferred technical solution, the video editing backbone network is built based on the text generation image model StableDiffusion, including CLIP text encoder, VAE autoencoder and UNet architecture;

[0024] The CLIP text encoder encodes the input text prompt information, the VAE autoencoder encodes the video into a latent variable space, and the UNet architecture extracts the local appearance of the video and the input text information through spatial and cross attention mechanisms;

[0025] Convert each convolutional layer and attention layer to a pseudo-3D layer by reshaping the frame axis to the batch axis;

[0026] The temporal attention module trained in Animatediff is integrated into the encoder and decoder of the UNet architecture to learn motion information across frames.

[0027] As a preferred technical solution, the multi-scale structural information feature map y output by the structural condition control network s The multi-scale appearance information feature map y output by the appearance condition control network a After addition, they are passed into the text generation image layer of the corresponding scale in the backbone network Unet architecture to generate the edited video.

[0028] As a preferred technical solution, in each temporal attention module in the video editing backbone network, the feature map is rearranged into a 3D tensor, and the self-attention mechanism is operated on the rearranged feature map:

[0029]

[0030] Where Q = W Qz,K=W K z,V=W V z are the three matrices obtained after projecting the rearranged feature maps.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The present invention proposes a versatile video editing framework that allows users to control the editing process by simultaneously inputting text prompts and detailed appearance information from images. In this framework, a structure-conditional control network captures the structural information of the original video, while an appearance-conditional control network extracts and propagates the appearance details from the edited image to the entire video. By combining these structural and appearance elements and leveraging various pre-trained text-to-image models, the gap between image editing and video editing can be effectively bridged, and videos with a wide range of structures, styles, and appearances can be created. The method of the present invention provides a flexible editing tool by coordinating appearance information and structural information. Users can combine various pre-trained personalized text-to-image generation models to edit videos according to specific needs and generate videos with a variety of styles, structures, and appearances.

[0033] 2) The present invention builds a video editing framework based on AnimateDiff, by reshaping the frame axis into a batch axis, converting each convolutional layer and attention layer into a pseudo 3D layer, thereby expanding the StableDiffusion model so that it can independently process each frame in the video, and integrating the temporal attention layer in the encoder and decoder of the backbone network UNet architecture to learn the motion information across frames, thereby promoting smooth dynamics of the video and enhancing the correlation between frames. The framework proposed by the present invention can easily integrate various personalized text-image generation models, enabling it to generate high-quality videos of various styles.

[0034] 3) The appearance conditional generation network proposed in the present invention can control the appearance attributes of each frame in the output video by editing the image of the first frame of the input video, thereby affecting the visual effect of the entire video. In this way, users can have fine-grained control over the appearance of the generated video to meet different visual needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 The flowchart of the controllable video editing method based on structure-appearance information fusion of the present invention;

[0036] Figure 2 A schematic diagram of the structure of a video editing framework based on structure-appearance information fusion according to the present invention;

[0037] Figure 3 It is a specific network structure diagram of the appearance condition control network of the present invention. DETAILED DESCRIPTION

[0038] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0039] Example 1

[0040] The present invention provides an innovative controllable video editing method based on structure-appearance information fusion, which can propagate the appearance information of a single frame edit to the entire video during the editing process. Specifically, the framework consists of three networks. The backbone network adopts a pre-trained text-to-image (T2I) diffusion model and is converted into a text-to-video (T2V) model through a temporal module. The structure condition control network encodes structural information using ControlNet. The appearance condition control network encodes user-defined appearance details for video generation. The network includes an appearance encoder and an appearance propagation module. The appearance encoder captures multi-scale features from the input image and propagates these features to the entire video through the appearance propagation module. Users can select various structural conditions, such as depth maps, line drawings, HED soft edges, and poses. For appearance control, users can use ControlNet or other online advanced image editing tools to edit an image and use the edited image as appearance condition information. In addition, various personalized style T2I models trained by DreamBooth or LoRA can be integrated into the framework to enhance the diversity of styles. These controls enable our framework to generate edited videos with diverse styles, structures, and appearances. Figure 1 As shown, a controllable video editing method based on structure-appearance information fusion proposed by the present invention comprises the following steps:

[0041] S1. Use the structural information extraction module to extract structural information (such as depth, posture, line drawing, edge line information, etc.) from the input video to be edited, and the obtained structural information is used as a structural control condition in the subsequent video editing process.

[0042] S2. Use the image editing module to edit the first frame of the input video, and the edited image is used as an appearance control condition in the subsequent video editing process.

[0043] S3. Build a structural conditional control network based on ControlNet, extract a multi-scale structural information feature map from the structural information obtained in step (S1), and provide structural information control for the subsequent video editing process.

[0044] S4. Build an appearance condition control network based on SparseCtrl, extract a multi-scale appearance information feature map from the appearance information obtained in step (S2), and provide appearance information control for the subsequent video editing process.

[0045] S5. Build the main framework of video editing based on AnimateDiff, fuse the multi-scale structural information feature map and the multi-scale appearance information feature map, and combine the input text information to generate the edited video.

[0046] In step S1, the extracted structural information includes depth, posture, line drawing, edge line information, etc.

[0047] In step S2, the image editing tools that can be used include online image editing tools, open source image editing tools, and the image editing tool StableDiffusion+ControlNet.

[0048] The backbone network of the video editing framework is built based on the text generation image model Stable Diffusion, integrating the CLIP text encoder, VAE autoencoder, and the main UNet architecture. The CLIP text encoder is used to encode text information, and the VAE autoencoder reduces the computational cost by encoding the video into the latent variable space. In the UNet architecture, the mutual correlation between local video frames and the input text information are introduced to the framework through the spatial and cross-attention mechanism. However, Stable Diffusion cannot be used directly for video generation, so in order to apply it to the field of video editing, the present invention builds a video editing main framework based on AnimateDiff. Specifically, by reshaping the frame axis into the batch axis, each convolutional layer and attention layer is converted into a pseudo-3D layer, thereby expanding the StableDiffusion model so that it can independently process each frame in the video. In addition, we integrate the time attention module trained in Animatediff into the encoder and decoder of the UNet architecture to learn the motion information across frames, thereby promoting the smooth dynamics of the video and enhancing the correlation between frames. The framework proposed in this invention can easily integrate various personalized text-image generation models, such as models trained with DreamBooth or LoRA, so that it can generate high-quality videos in various styles, such as animation, Van Gogh, watercolor, oil painting style, etc.

[0049] For the Unet architecture and each temporal attention layer in AnimateDiff, let z be of shape We first rearrange the feature map into a 3D tensor Subsequently, the rearranged feature map undergoes a series of operations mainly involving the self-attention mechanism. Specifically, the present invention adopts the following attention mechanism:

[0050]

[0051] Where Q = W Q z,K=W K z,V=W V z are the three matrices obtained by projecting the rearranged feature maps. This attention mechanism effectively captures the interdependence between different frames and enhances the model's ability to understand the temporal relationship in the data.

[0052] The structural conditional control network uses ControlNet (a pre-trained model for controlled image generation) to encode different structural information frame by frame. The type of ControlNet corresponds to the type of structural information selected. Each type of structural information corresponds to a ControlNet. Whenever a type of structural information is input, the corresponding ControlNet is used. For example, if the deep structural information is extracted in step S1, then the structural conditional control network in step S3 uses the ControlNet in which the deep information corresponds to the deep structural information. Subsequently, the encoded information is seamlessly integrated into the decoder of the backbone network of the video editing framework in step S5 to control the video generation process. F1 represents the encoder block of the backbone network, F2 represents the encoder part of the structural conditional control network ControNet corresponding to the backbone network encoder F1, and Z1 and Z2 represent two zero-initialized convolutional layers. The entire process of the structural conditional control network can be expressed as follows:

[0053] y s =F1(x)+α s *Z1(F2(x+Z2(c s )))

[0054] Among them, α s is a hyperparameter that controls the strength of the network under the structural condition, and is set to 0.5 by default. The video in the latent variable space, c s represents the structure representation extracted from the input video, y s Represents the output feature map containing structural information.

[0055] like Figure 3 As shown in the figure, the first frame of the edited image will be spliced ​​with a completely black frame, and there will be a conditional mask of frame T (the first frame is all 1, and the following frames are all 0, indicating that only the first frame is a meaningful edited image). The two are concatenated together in the channel dimension and then input into the appearance condition control network.

[0056] The appearance conditional control network extracts appearance features from the edited frames and propagates them to the entire video, providing users with fine-grained control over visual aesthetics. The appearance conditional control network consists of two main components: the appearance encoder and the appearance information propagation module. The output of the appearance conditional control network is a multi-scale feature map, which is then directly added to different layers of the main Unet network according to shape correspondence.

[0057] Appearance encoder: The network structure of this component is similar to the frame-by-frame encoder in ControlNet. On the basis of ControlNet, each convolutional layer and attention layer is expanded from 2D to pseudo-3D layer to expand its architecture so that it can be applied in the field of video editing. The input I of the appearance encoder has multiple frames, I0 (the first frame) is the appearance control picture edited by the image editing tool in step S2, and the remaining frames are feature maps with zero values. The role of the appearance encoder is to inject the appearance condition information of I0 into the video editing process, thereby controlling the appearance of the output edited video.

[0058] Appearance Information Propagation Module: This component plays a vital role in maintaining the consistency of appearance information across frames of the edited video. The present invention incorporates an appearance propagation module after each appearance encoder. Specifically, the appearance information propagation module utilizes a temporal attention layer to propagate the appearance condition information extracted by the appearance encoder to each frame of the video, aiming to improve the visual consistency and coherence of the output video.

[0059] Let F3 be an encoder module of the backbone network of the video editing framework, F4 be an encoder module of the appearance condition control network (including the appearance encoder and the appearance propagation module), and Z3 and Z4 be two zero-initialized convolutional layers. The processing of the appearance condition control network is as follows:

[0060] y a =F3(x)+α a *Z3(F4(Z4(Concat(I,M))))

[0061] Among them, α a It is a hyperparameter that controls the appearance condition and the network strength. The default setting is 1; x represents the shape y is a video in the latent variable space of y; I is a multi-frame input containing edited frames and several zero maps; M is a binary conditional mask containing the frames indicating edited images. a represents the feature map with appearance information, which will be combined with y in the structural conditional control network s After addition, it is passed to the corresponding T2I layer of the main Unet network to generate the edited video, so that both structure and appearance information can be used in video editing.

[0062] In summary, the method proposed in the present invention aims to make full use of the advantages of combining the structural information and appearance information of the video, overcome the limitations of existing video editing methods in accurately controlling the appearance of the output video, and realize efficient and flexible controllable video editing. First, ControlNet is introduced as a structural conditional control network to extract and inject various structural information from the input video, including but not limited to contours, edges, posture information, etc. These structural information provide a solid foundation for video editing, ensuring that the edited video retains the structural characteristics of the original video. Subsequently, the present invention also introduces an appearance conditional control network, which is used to combine an image edited by the user as the appearance control information in the video editing process. The network can control the appearance attributes of each frame in the output video through an edited picture, thereby affecting the visual effect of the entire video. In this way, the user can perform fine-grained control on the appearance of the generated video to meet different visual needs.

[0063] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A controllable video editing method based on structure-appearance information fusion, characterized in that: The method controls and edits the video to be edited based on the text prompt input by the user, and the specific steps include: Extract structural information from the video to be edited and input it into the structural condition control network based on ControlNet to obtain a multi-scale structural information feature map; Edit the first frame of the video to be edited and input it into the appearance condition control network based on SparseCtrl to obtain a multi-scale appearance information feature map; The multi-scale structural information feature map and the multi-scale appearance information feature map are fused and passed into the video editing backbone network based on AnimateDiff. The video editing backbone network generates an edited video based on the input text information.

2. A controllable video editing method based on structure-appearance information fusion according to claim 1, characterized in that: The structure condition control network uses ControlNet to encode different structure information frame by frame. The encoded structure information is integrated into the decoder of the video editing backbone network to control the video generation process. The whole process of the structure condition control network is expressed as follows: y s =F1(x)+α s *Z1(F2(x+Z2(c s ))) Among them, F1 represents the encoder block of the backbone network; F2 represents the encoder part of the structural conditional control network ControNet corresponding to the backbone network encoder F1, Z1 and Z2 represent two zero-initialized convolutional layers; α s is a hyperparameter that controls the structural conditions and network strength; x represents the shape The video in the latent variable space of c s represents the structure representation extracted from the input video; y s Represents the output feature map containing structural information.

3. The controllable video editing method based on structure-appearance information fusion according to claim 1, characterized in that: The structural information includes depth, posture, line drawing and edge lines.

4. The controllable video editing method based on structure-appearance information fusion according to claim 3, characterized in that: The structural condition control network constructs a ControlNet corresponding to each type of structural information; whenever a type of structural information is input, the corresponding ControlNet is used to extract a multi-scale structural information feature map.

5. The controllable video editing method based on structure-appearance information fusion according to claim 1, characterized in that: The method uses an image editing tool to edit the first frame of a video to be edited to obtain an appearance control picture; a frame with all zeros is spliced ​​after the appearance control picture, and is connected with a conditional mask with all 1s in the first frame and all 0s in subsequent frames in the channel dimension, and then input into an appearance conditional control network.

6. The controllable video editing method based on structure-appearance information fusion according to claim 5, characterized in that: The appearance condition control network includes an appearance encoder and an appearance information propagation module; The appearance encoder expands each convolution layer and attention layer from 2D to pseudo 3D layer based on ControlNet, and extracts appearance condition information from the appearance control picture; An appearance propagation module is incorporated after each appearance encoder, and the appearance information propagation module uses a temporal attention layer to propagate the appearance condition information extracted by the appearance encoder to each frame of the video.

7. The controllable video editing method based on structure-appearance information fusion according to claim 6, characterized in that: The processing process of the appearance condition control network is as follows: y a =F3(x)+α a *Z3(F4(Z4(Concat(I,M)))) Among them, y a represents a feature map with appearance information; F3 is an encoder module of the video editing backbone network, F4 is an encoder module of the appearance condition control network, Z3 and Z4 are two zero-initialized convolutional layers; α a is a hyperparameter controlling the strength of the appearance conditional control network1; x represents the video in the latent variable space; I is a multi-frame input containing edited frames and several zero maps; M contains a binary conditional mask indicating the edited image frame.

8. The controllable video editing method based on structure-appearance information fusion according to claim 1, characterized in that: The video editing backbone network is built based on the text-generated image model Stable Diffusion, which includes CLIP text encoder, VAE autoencoder and UNet architecture; The CLIP text encoder encodes the input text prompt information, the VAE autoencoder encodes the video into a latent variable space, and the UNet architecture extracts the local appearance of the video and the input text information through spatial and cross attention mechanisms; Convert each convolutional layer and attention layer to a pseudo-3D layer by reshaping the frame axis to the batch axis; The temporal attention module trained in Animatediff is integrated into the encoder and decoder of the UNet architecture to learn motion information across frames.

9. The controllable video editing method based on structure-appearance information fusion according to claim 8, characterized in that: The structural condition controls the multi-scale structural information feature map y output by the network s The multi-scale appearance information feature map y output by the appearance condition control network a After addition, they are passed into the text generation image layer of the corresponding scale in the backbone network Unet architecture to generate the edited video.

10. The controllable video editing method based on structure-appearance information fusion according to claim 8, characterized in that: In each temporal attention module in the video editing backbone network, the feature map is rearranged into a 3D tensor, and the self-attention mechanism is operated on the rearranged feature map: Where Q = W Q z,K=W K z,V=W V z are the three matrices obtained after projecting the rearranged feature maps.