Video editing method based on improved pre-trained diffusion model
Patent Information
- Application Number
- CN202311806284.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-26
AI Technical Summary
该方法使用空间自注意力机制,在视频编辑时引入了空间信息,一定程度上提升了编辑视频的帧一致性,且引入了参考图像信息,对视频的主角进行更直观的描述,一定程度上提升了编辑视频的美观性;但其存在的缺陷在于计算空间自注意力时,每一个图像帧只与第一个图像帧和前一个图像帧进行耦合,在视频主角动作幅度较大时,无法提供足够的空间指导,导致编辑视频的帧一致性仍然较差,且在去噪的过程中,损害了一些高频特征信息,导致编辑视频的美观性仍然较差
[0020](1)本发明通过在预训练扩散模型所包含的去噪采样模块中每个卷积模块与交叉注意力模块之间加载空间自注意力模块,同时在去噪采样模块解码器端上采样模块和每个注意力上采样模块输入端加载频率优化模块,得到改进的去噪采样模块,在对其进行训练以及获取源视频编辑结果的过程中,空间自注意力模块能够获取所有图像帧的空间特征,在视频主角动作幅度较大时,能为视频的编辑过程提供丰富的空间信息指导,与现有技术相比,有效提高了编辑视频的帧一致性。
Smart Images

Figure CN117834987B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video generation technology and relates to a video editing method, specifically a video editing method based on an improved pre-trained diffusion model, which can be applied to video style conversion, video content and background replacement. Background Technology
[0002] Traditional video editing primarily uses Generative Adversarial Networks (GANs) to decouple features in the vector space, allowing changes to one feature to avoid affecting another. However, GANs suffer from training instability, vanishing gradients, and mode collapse. Current video editing typically uses diffusion models to add noise to the source video, introducing editing prompts during the denoising and sampling process. However, when using only editing prompts to guide the denoising and sampling process, the trained text encoder has a limited capacity for encoding text content. It cannot describe specific individuals in text, or use text outside the scope of its descriptive capabilities. To reduce the impact of single text descriptions on the frame consistency and aesthetics of edited videos, researchers have improved the conditions for denoising sampling in guided diffusion models. For example, in the paper "Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts" published on May 15, 2023, Yuyang Zhao disclosed a video editing method based on text image guidance. This method uses multiple expert models to parse the source video, reference image, and editing prompts, and combines a text image-based video generation model with a mask-based denoising sampling algorithm to enable editing prompts and reference images to jointly edit the source video. This method uses a spatial self-attention mechanism to introduce spatial information during video editing, which improves the frame consistency of the edited video to some extent. It also introduces reference image information to provide a more intuitive description of the main character in the video, thus improving the aesthetics of the edited video to some extent. However, its drawback is that when calculating spatial self-attention, each image frame is only coupled with the first image frame and the previous image frame. When the main character's movements are large, it cannot provide sufficient spatial guidance, resulting in poor frame consistency in the edited video. Furthermore, during the denoising process, some high-frequency feature information is damaged, resulting in poor aesthetics in the edited video. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the existing technology and propose a video editing method based on an improved pre-trained diffusion model, which aims to improve the frame consistency and aesthetics of edited videos.
[0004] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:
[0005] (1) Constructing training samples and inference samples:
[0006] Construct a source video consisting of N image frames. and its description and reference image The training samples R1, and the source video consisting of N image frames. and its editing prompts and reference image The inference samples R², N≥8;
[0007] (2) Improve the denoising sampling module in the pre-trained diffusion model:
[0008] In the denoising sampling module included in the pre-trained diffusion model, a spatial self-attention module is loaded between each convolutional module and the cross-attention module. Simultaneously, a frequency optimization module is loaded at the input of the upsampling module at the decoder end of the denoising sampling module and at the input of each attention upsampling module, resulting in an improved denoising sampling module. θ The spatial self-attention module is a residual structure, including a cascaded normalization layer, a frame feature stitching module, and an attention calculation module; the frequency optimization module includes a cascaded fast Fourier transform module (FFT), a low-frequency feature reduction module, an inverse fast Fourier transform module (IFFT), and a feature stitching module, and the input of the feature stitching module is also loaded with a feature amplification module.
[0009] (3) Initialize parameters:
[0010] The initial number of iterations is s, the maximum number of iterations is S, S≥200, and the weight parameter of the spatial self-attention module in the s-th iteration is θ. s And let s = 1;
[0011] (4) Extract features from training samples:
[0012] The source video is extracted based on time step t by the cascaded encoder ε and forward noise module in the pre-trained diffusion model. Each image frame in noise characteristics Extracting descriptions using CLIP text-image encoder Feature I prompt and reference image Feature I img and will I prompt and I img The training sample feature set is composed of;
[0013] (5) Train the improved denoising sampling module:
[0014] The training sample feature set and time step t are used as inputs to the improved denoising sampling module to process the noise feature x. vnt The noise added to the sample is estimated to obtain the estimated noise.
[0015] (6) Obtain the trained improved denoising sampling module:
[0016] Through noise For the weight parameter θ s The module was updated to obtain the improved denoising sampling module for this iteration. And determine whether s = S holds true. If so, obtain the trained improved denoising sampling module ∈ θ * Otherwise, let s = s + 1. And proceed to step (4);
[0017] (7) Obtain the source video editing result based on the video editing model H of the improved denoising sampling module:
[0018] A video editing model H based on an improved denoising sampling module is constructed, and inference samples R2 and time step t′ are used as inputs to the video editing model H for inference to obtain the source video. Editing videos
[0019] Compared with the prior art, the present invention has the following advantages:
[0020] (1) This invention improves the denoising sampling module by loading a spatial self-attention module between each convolutional module and the cross-attention module in the denoising sampling module included in the pre-trained diffusion model, and by loading a frequency optimization module at the input end of the upsampling module at the decoder end of the denoising sampling module and at the input end of each attention upsampling module. During the training of the module and the acquisition of the source video editing results, the spatial self-attention module can acquire the spatial features of all image frames. When the main character in the video has a large range of motion, it can provide rich spatial information guidance for the video editing process. Compared with the prior art, it effectively improves the frame consistency of the edited video.
[0021] (2) The improved denoising sampling module used in this invention includes a frequency optimization module to obtain an improved denoising sampling module. During the training of the improved denoising sampling module and the acquisition of source video editing results, the frequency optimization module reduces the low-frequency features introduced by the jump connection and correspondingly amplifies the high-frequency features, making the texture and edge details of the generated video better. Compared with the prior art, it effectively improves the aesthetics of the edited video. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the implementation of the present invention;
[0023] Figure 2 This is a schematic diagram of the structure of the pre-trained diffusion model used in this invention;
[0024] Figure 3 This is a schematic diagram of the denoising sampling module in the pre-trained diffusion model used in this embodiment of the invention;
[0025] Figure 4 This is a schematic diagram of the composite module in the pre-trained diffusion model used in this invention;
[0026] Figure 5 This is a schematic diagram of the spatial self-attention module in the improved denoising sampling module used in this invention;
[0027] Figure 6 This is a schematic diagram of the frequency optimization module in the improved noise reduction sampling module used in this invention;
[0028] Figure 7 This is a schematic diagram of the video editing model based on the improved denoising sampling module used in this invention. Detailed Implementation
[0029] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0030] Reference Figure 1 The present invention includes the following steps:
[0031] Step 1) Construct training samples and inference samples:
[0032] (1a) Extract the source video consisting of N image frames. Description And randomly select any frame of the image as the reference image. Then and Combined into training samples R1, where The nth image frame is N≥8;
[0033] (1b) Initialize the source video The editing prompt is: And select an image from the network. As The reference image, then and Combined into inference sample R2. The selected reference image should be roughly similar in appearance to the main character in the source video; in this embodiment, N = 8.
[0034] Step 2) Improve the denoising sampling module in the pre-trained diffusion model:
[0035] Pre-trained diffusion models such as Figure 2 As shown, it includes a cascaded encoder ε, a forward noise-adding module, a denoising sampling module with a UNet network structure, and a decoder. A DDIM inversion module is also loaded between the encoder ε and the denoising sampling module, and a CLIP text image encoder is loaded at the input of the denoising sampling module, wherein:
[0036] The structure of the UNet network is as follows: Figure 3 As shown, the encoder consists of an encoder consisting of M cascaded attention downsampling modules and a downsampling module; the decoder consists of a cascaded upsampling module and M attention upsampling modules; and an attention module connects the encoder and decoder. The m-th attention downsampling module is connected to the M-(m-1)-th attention upsampling module. Each attention downsampling module includes multiple cascaded composite modules and a downsampling block. Each downsampling module includes multiple cascaded residual modules and a downsampling block. Each upsampling module includes multiple cascaded residual modules and an upsampling block. Each attention upsampling module includes multiple cascaded composite modules and an upsampling block. Each attention module includes a cascaded composite module and a residual module. The structure of the composite module is as follows: Figure 4 As shown, it includes cascaded residual modules, convolutional modules, cross-attention modules, feedforward modules, and convolutional layers. The convolutional modules include cascaded group normalization layers and convolutional layers. The CLIP text image encoder is a pre-trained expert model, where M ≥ 3. In this embodiment, M = 3.
[0037] In the denoising sampling module included in the pre-trained diffusion model, a spatial self-attention module is loaded between each convolutional module and the cross-attention module. Simultaneously, a frequency optimization module is loaded at the input of the upsampling module at the decoder end of the denoising sampling module and at the input of each attention upsampling module, resulting in an improved denoising sampling module. θ The structure of the spatial self-attention module is as follows: Figure 5 As shown, it includes a cascaded normalization layer, a frame feature concatenation module, and an attention calculation module; the structure of the frequency optimization module is as follows. Figure 6 As shown, it includes a cascaded Fast Fourier Transform (FFT) module, a low-frequency feature reduction module, an Inverse Fast Fourier Transform (IFFT) module, and a feature concatenation module, and the input of the feature concatenation module is also loaded with a feature amplification module.
[0038] Step 3) Initialize parameters:
[0039] The initial number of iterations is s, the maximum number of iterations is S, S≥200, and the weight parameter of the spatial self-attention module in the s-th iteration is θ. s And let s = 1; in this embodiment, S = 200;
[0040] Step 4) Extract features from training samples:
[0041] (4a) Encoder ε for each image frame Encode, Transforming from the pixel domain to the latent space yields... Potential characteristics The forward noise module adjusts the noise level according to time step t. Adding Gaussian noise ∈, we get noise characteristics
[0042]
[0043]
[0044] in, Indicates a total of t steps α i The cumulative product, α i It is the hyperparameter for adding noise in the i-th step;
[0045] Time step t refers to the number of times the forward noise-adding module in the pre-trained diffusion model adds noise, where t∈[1,T1], obtained by uniform sampling in [1,T1], and T1 represents the maximum number of noise-adding times, T1≥1000; in this embodiment, T1=1000;
[0046] (4b) Extracting descriptions using CLIP text-image encoder Feature I prompt and reference image Feature I img and will I prompt and I img The training sample feature set is composed of;
[0047] Step 5) Train the improved denoising sampling module:
[0048] The training sample feature set and time step t are used as inputs to the improved denoising sampling module to process noise features. The noise added to the image is estimated, and the steps are as follows:
[0049] (5a) Improved denoising sampling module ∈ θ The m-th attention downsampling module at the encoder end Feature extraction is performed to obtain coupled features. Again Perform downsampling to obtain downsampled features.
[0050] (5a1) When m=1, the improved denoising sampling module ∈ θThe residual module of the m-th attention downsampling module at the encoder encodes time step t to obtain the feature I of time step t. t and will I img and I t The fusion process incorporates temporal and reference image information into the noise features to obtain fused features. The convolutional module then normalizes and extracts features from the fused features to obtain normalized features.
[0051] (5a2) Spatial self-attention mechanism can accurately analyze the spatial information of images, helping the model understand the content and spatial structure of visual images. By enhancing the feature representation of key regions, it helps to strengthen specific target regions of interest while weakening irrelevant background regions. To fully utilize the spatial information of each image frame in the source video, and to provide rich spatial information guidance for the video editing process when the main character's movements are large, the spatial self-attention module in the composite module will... The concatenated features of all N image frames To perform coupling, we obtain coupling characteristics.
[0052]
[0053]
[0054]
[0055]
[0056] Q1 represents The query, W K W V d1 represents the parameter matrix for calculating key K1 and value V1, respectively, and d1 represents the length of K1;
[0057] (5a3) To incorporate textual information for editing the source video, the cross-attention module in the composite module will... and I prompt By coupling, new coupling characteristics are obtained.
[0058]
[0059]
[0060] K2 = W K I prompt
[0061] V2 = W V I prompt
[0062] Q1 represents In the query, d2 represents the length of K2;
[0063] (5a4) Feedforward module Feature extraction is performed, the convolutional layer extracts features again, and the downsampling block downsamples the convolutional features to obtain downsampled features;
[0064] (5a5) When m = 2, 3, the m-th attention downsampling module extracts features from the downsampling features output by the (m-1)-th attention downsampling module to obtain the downsampling features;
[0065] (5b) Improved denoising sampling module ∈ θ The residual module of the encoder's downsampling module outputs the downsampled features of the third attention downsampling module. I img and I t Feature fusion is performed, and the downsampling block downsamples the fused features to obtain new downsampled features.
[0066] (5c) Spatial self-attention module and cross-attention module of the attention module Perform coupling operations to obtain hybrid features.
[0067] (5d) Frequency domain signal processing can decompose a signal into different frequency components, which can better characterize the signal's features. U-Net impairs high-frequency details during denoising, while its skip connections directly forward features from earlier layers of the encoder to the decoder; these features primarily constitute high-frequency information. Fast Fourier Transform (FFT) provides accurate frequency component measurements, allowing for better adjustment of the high and low frequency components of the source video features in the frequency domain. By using FFT to centralize the low-frequency information of skip features, the low-frequency information of the skip features is reduced, while the high-frequency information is correspondingly increased, resulting in better image texture and edge details. Improved denoising sampling module ∈ θ The frequency optimization module of the upsampling module at the decoder end uses right Frequency optimization is performed to obtain frequency optimization features.
[0068]
[0069]
[0070]
[0071] Where γ(r) represents The parameters of the feature whose distance from the feature center point is r in the frequency domain feature. express The parameters of the c-th channel are: ⊙ represents element-wise multiplication, thrd represents the threshold, s represents values less than 1, and C represents the feature. The number of channels, where b represents a value greater than 1;
[0072] Residual module pairs Feature fusion is performed, and the upsampling block upsamples the fused features to obtain upsampled features.
[0073] (5e) Improved denoising sampling module ∈ θ The M-(m-1)th attention upsampling module at the decoder end uses right After frequency optimization, feature extraction is performed to obtain denoised features. according to and The difference is used to obtain the estimated noise.
[0074] (5e1) When M-(m-1)=1, the improved denoising sampling module ∈ θ The frequency optimization module of the M-(m-1)th attention upsampling module at the decoder end uses right Frequency optimization is performed. The composite module extracts features from the optimized features, and the upsampling block upsamples the extracted features to obtain optimized upsampled features.
[0075] (5e2) When M-(m-1)=2,3, the improved denoising sampling module ∈ θ The frequency optimization module of the M-(m-1)th attention upsampling module at the decoder uses the output of the m-th attention downsampling module at the encoder to perform frequency optimization on the optimized upsampling features output by the Mm-th attention upsampling module. The composite module extracts features from the optimized features, and the upsampling block upsamples the extracted features to obtain the denoised features output by the third attention upsampling module. according to and The difference is obtained from the denoising sampling module ∈ θ Estimated noise in:
[0076]
[0077]
[0078] Where z is the noise parameter.
[0079] Step 6) Obtain the trained improved denoising sampling module:
[0080] Through noise For the weight parameter θ s The module was updated to obtain the improved denoising sampling module for this iteration. And determine whether s = S holds true. If so, obtain the trained improved denoising sampling module ∈ θ * Otherwise, resample the time step t, letting s = s + 1. And perform step 4), where θ s The update formula is:
[0081]
[0082]
[0083] Where, θ s 'represents θ s The update result, l r Indicates the learning rate. L represents the differentiation operation. s This represents the loss value of the video editing model calculated using the mean squared error loss function. The expression represents the expectation operation, and |||·||| represents the L1 norm operation.
[0084] Step 7) Obtain the source video editing result based on the video editing model H of the improved denoising sampling module:
[0085] Construct a video editing model H based on an improved denoising sampling module, the structure of which is as follows: Figure 7 As shown, it includes a cascaded pre-trained encoder ε, a pre-trained DDIM inversion module, a feature fusion module, and an improved denoising sampling module ∈θ. * and pre-trained decoder The feature fusion module's input also includes a parallel-connected mask segmentation unit and a CLIP text image encoder, improving the denoising sampling module. θ * The input terminal is also loaded with a cascaded control signal extraction unit and a ControlNet unit; the inference sample R2 and time step t′ are used as inputs to the video editing model H for inference, where t′∈[1,T2], T2≥25, and in this embodiment T2=50. The implementation steps are as follows:
[0086] (7a) Pre-trained encoder ε pairs with source video Each image frame in Feature extraction is performed to obtain Potential characteristics Pre-trained DDIM inversion module Inversion is performed to obtain noise characteristics. Provides structural guidance for the denoising sampling process; the mask segmentation unit acquires the source video. Mask in Separating the main subject from the background in the source video; CLIP text-image encoder extracts editing prompts. Features and reference image Features Among them, the mask segmentation unit is a pre-trained expert model;
[0087] (7b) In order to concentrate the reference image features on the main character part of the source video and the editing prompt word features on the background part of the source video, so as to realize the replacement of the main character of the video and the control of the background part through the editing prompt word, the feature fusion module uses a mask. in right and The fusion process is performed to obtain the fused noise characteristics.
[0088]
[0089] (7c) To supplement the depth and pose information of the edited video, the control signal extraction unit extracts... Depth map M depth and attitude diagram M pose ControlNet unit extracts M depth M pose Feature I depth I pose Improved denoising sampling module ∈ θ * The residual modules of the encoder and attention module are related to I. depth I pose and By fusing, conditional features are obtained. And on After performing T2-step denoising sampling, the feature x is obtained. n0 ",where the control signal extraction unit and the ControlNet unit are both pre-trained expert models;
[0090] (7d) Pre-trained decoder For x n0 "Decode x" n0 "Transforming from the latent space to the pixel domain yields the source video." Editing videos
[0091] The technical effects of the present invention will be further explained below with reference to simulation results:
[0092] 1. Experimental conditions and contents:
[0093] The experimental hardware platform consisted of an Intel(R) Xeon(R) CPU with a clock speed of 2.40GHz, 125GB of RAM, and an NVIDIA A100 graphics card with 80GB of video memory. The software environment included Ubuntu 22.04.1 operating system, Python 3.9.16, and PyTorch 2.0.0. The dataset used was DAVIS.
[0094] The frame consistency and aesthetics of this invention were compared with those of the existing Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts through simulation, and the results are shown in Table 1. Frame consistency refers to the average cosine similarity of the embeddings of all video frame pairs in the output edited video. Aesthetics were evaluated using two metrics: CLIP-T and DINO. CLIP-T refers to the average cosine similarity between the embeddings of all frames in the output edited video and the corresponding text editing prompt CLIP embeddings. DINO refers to the average cosine similarity between the embeddings of all frames in the output edited video and the corresponding reference image embeddings.
[0095] 2. Analysis of experimental results:
[0096] Table 1
[0097] Existing technology 30.98 93.26 30.78 This invention 32.09 94.42 43.82
[0098] Referring to Table 1, the present invention outperforms the Make-A-Protagonist method by 1.11%, 1.16%, and 13.04% in CLIP-T, frame consistency, and DINO metrics on the DAVIS dataset, respectively. The above experimental results demonstrate that the present invention effectively improves the frame consistency and aesthetics of edited videos compared to existing technologies.
Claims
1. A video editing method based on an improved pre-trained diffusion model, characterized in that, Includes the following steps: (1) Constructing training samples and inference samples: Build including The source video of one image frame and its description and reference image training samples and including The source video of one image frame and its editing prompts and reference image Reasoning samples , ; (2) Improve the denoising sampling module in the pre-trained diffusion model: In the denoising sampling module included in the pre-trained diffusion model, a spatial self-attention module is loaded between each convolutional module and the cross-attention module. Simultaneously, a frequency optimization module is loaded at the input of the upsampling module at the decoder end of the denoising sampling module and at the input of each attention upsampling module, resulting in an improved denoising sampling module. The spatial self-attention module is a residual structure, comprising a cascaded normalization layer, a frame feature concatenation module, and an attention calculation module. The frequency optimization module includes a cascaded Fast Fourier Transform (FFT) module, a low-frequency feature reduction module, an Inverse Fast Fourier Transform (IFFT) module, and a feature concatenation module. Furthermore, the input of the feature concatenation module is loaded with a feature amplification module. Pre-trained diffusion model, including loaded cascaded encoders A forward noise-adding module, a denoising sampling module with a UNet network structure, and a decoder. and encoder A DDIM inversion module is also loaded between the denoising and sampling modules, and a CLIP text-image encoder is also loaded at the input of the denoising and sampling module; the UNet network consists of cascaded... The encoder end consists of an attention downsampling module and a downsampling module, and is composed of cascaded upsampling modules and The decoder consists of several attention upsampling modules, and the attention module connects the encoder and decoder. The attention downsampling module and the first Attention upsampling modules are connected via skip connections; attention downsampling modules include multiple cascaded composite modules and a downsampling block; downsampling modules include multiple cascaded residual modules and a downsampling block; upsampling modules include multiple cascaded residual modules and an upsampling block; attention upsampling modules include multiple cascaded composite modules and an upsampling block; attention modules include cascaded composite modules and a residual module; composite modules include cascaded residual modules, convolutional modules, cross-attention modules, feedforward modules, and convolutional layers; convolutional modules include cascaded group normalization layers and convolutional layers, wherein... ; (3) Initialize parameters: Initialize the number of iterations to The maximum number of iterations is , , No. The weight parameters of the next iteration spatial self-attention module are: and order ; (4) Extracting features from training samples: The encoder cascaded in the pre-trained diffusion model and the forward noise module according to time steps Extract source video Each image frame in noise characteristics Extract description using CLIP text-image encoder Features and reference image Features and will , and The training sample feature set is composed of; (5) Train the improved denoising sampling module: The training sample feature set and time step As input to the improved denoising sampling module, the noise characteristics are... The noise added to the sample is estimated to obtain the estimated noise. ; (6) Obtain the trained improved denoising sampling module: Through noise For weight parameters The module was updated to obtain the improved denoising sampling module for this iteration. and judge If true, then obtain the trained improved denoising sampling module. Otherwise, re-determine the time step. At the same time, , and perform step (4); (7) Video editing model based on improved denoising sampling module Get the results of editing the source video: Constructing a video editing model based on an improved denoising sampling module and inference samples and time step As a video editing model The input is used for reasoning to obtain the source video. Editing videos .
2. The method according to claim 1, characterized in that, The steps for constructing training samples and inference samples as described in step (1) are as follows: (1a) Extraction includes The source video of one image frame Description And randomly select any frame of the image as the reference image. Then , and Combined into training samples ,in The One image frame is ; (1b) Initialize the source video The editing prompt is: And select an image from the network. As The reference image, then , and Combined into reasoning samples .
3. The method according to claim 1, characterized in that, The time step described in step (4) , refers to the number of times the forward noise-adding module in the pre-trained diffusion model adds noise, where, , Indicates the maximum number of noise additions. .
4. The method according to claim 1, characterized in that, The extraction of source video described in step (4) Each image frame in noise characteristics The implementation steps are as follows: encoder For each image frame Encode to obtain Potential characteristics The forward noise addition module is based on the time step. right Add Gaussian noise ,get noise characteristics : ; ; in, Indicates total step The cumulative product, It is the first Hyperparameters for adding noise.
5. The method according to claim 1, characterized in that, The improved denoising sampling module described in step (5) The training process involves the following steps: (5a) Improved denoising sampling module The encoder end Each attention downsampling module Feature extraction is performed to obtain coupled features. And then Perform downsampling to obtain downsampled features. ,in: ; ; ; ; ; ; ; ; in, express , and time step Features The fusion characteristics This represents the features after the spatial self-attention module is coupled. , They represent , The query, , These represent the calculation keys. , Sum , The parameter matrix, , express , Length, Indicates all The splicing features of image frame features; (5b) Improved denoising sampling module The downsampling module at the encoder end Perform downsampling to obtain downsampled features. ; (5c) Attention module Perform coupling operations to obtain hybrid features. ; (5d) Improved denoising sampling module The upsampling module at the decoder end uses right Frequency optimization is performed to obtain frequency optimization features. ,right Feature extraction is performed to obtain upsampled features. ,in: ; ; ; in, express The distance from the feature center point in the frequency domain features is The parameters of the features, express No. Parameters for each channel, Indicates element-wise multiplication. Indicates the threshold. This represents a value less than 1. Representation of features The number of channels, This represents a value greater than 1; (5e) Improved denoising sampling module The decoder end An attention upsampling module uses right After frequency optimization, feature extraction is performed to obtain denoised features. ; and according to and The difference is used to obtain the estimated noise. ,in: ; ; in, It is a noise parameter.
6. The method according to claim 5, characterized in that, The step (6) described above involves passing through noise. For weight parameters The update is performed using the following formula: ; ; in, express The update results Indicates the learning rate. This represents the differentiation operation. This represents the loss value of the video editing model calculated using the mean squared error loss function. This indicates the operation of finding the expected value. This indicates the operation of calculating the L1 norm.
7. The method according to claim 6, characterized in that, The video editing model based on the improved denoising sampling module described in step (7) Including cascaded pre-trained encoders Pre-trained DDIM inversion module, feature fusion module, and improved denoising sampling module and pre-trained decoder The feature fusion module's input also includes a parallel-connected mask segmentation unit and a CLIP text image encoder, improving the denoising sampling module. The input terminal is also loaded with a cascaded control signal extraction unit and a ControlNet unit.
8. The method according to claim 7, characterized in that, The reasoning sample described in step (7) and time step As a video editing model The inference process is performed based on the input, and the steps are as follows: (7a) Pre-trained encoder For the source video Each image frame in Feature extraction is performed to obtain Potential characteristics The pre-trained DDIM inversion module is used for... Inversion is performed to obtain noise characteristics. The mask segmentation unit obtains the source video. mask ; CLIP Text Image Encoder Extracts Editing Prompts Features and reference image Features ,in, , ; (7b) The feature fusion module uses a mask. right , and The fusion process is performed to obtain the fused noise characteristics. : ; (7c) Control signal extraction unit extracts Depth map and attitude diagram ControlNet unit extraction , Features , Improved noise reduction sampling module right , and After fusion, denoising sampling is performed to obtain the denoised sampling features. ; (7d) Pre-trained decoder right Decode to obtain the source video Editing videos .
Citation Information
Patent Citations
Image multistage denoising method based on deep learning
CN111598804A
Self-attention mechanism-based behavior recognition method
WO2022083335A1