Video frame insertion method of generative video diffusion model based on event guidance

Through the event-guided generative video diffusion model, combined with the multimodal motion condition generator and selective time-domain self-attention fine-tuning strategy, the existing video interpolation technology has solved the insufficient performance problems of large motion, complex lighting changes and low-texture areas, and achieved high-quality video interpolation effect.

CN120475196APending Publication Date: 2025-08-12SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510352485.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing video interpolation technology has poor performance when dealing with large motion, complex lighting changes and low-texture areas. Traditional methods have problems such as inaccurate optical flow estimation, insufficient utilization of event data, and lack of timing guidance for diffusion models.

Method used

Using an event-guided generative video diffusion model, an event-guided generative video diffusion network is constructed, including a multimodal motion condition generator, a spatiotemporal attention condition U-type network and a guided embedded feature aggregation module. Combined with a two-stage training strategy, the high-temporal resolution information of the event stream and the powerful generation ability of the diffusion model are integrated.

Benefits of technology

High-quality video interpolation in large motion, complex lighting changes and low-texture areas are achieved, which improves the generalization ability of the model and the quality of the generation details, reduces the difficulty of training, and improves the interpolation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475196A_ABST
    Figure CN120475196A_ABST
Patent Text Reader

Abstract

The invention relates to a video frame insertion method of a generative video diffusion model based on event guidance, and the method comprises the following steps: constructing an event guidance generative video diffusion network, the event-guided generative video diffusion network comprises a condition generator MMCG, a pre-trained VAE encoder / decoder, a space-time attention condition U-type network Unet and a guided embedded feature aggregation module GEA, and the condition generator MMCG comprises a region of interest ROI selector, a voxel feature extractor VFE, a multi-modal feature fusion module MMF and a three-dimensional residual learning module 3DVR; and adopting a two-stage training strategy: inputting the video data and the event stream data into the guide generation type video diffusion model to obtain an interpolation frame. Compared with the prior art, the method has the advantage of overcoming the defect that the existing video frame insertion technology is poor in performance when processing large-amplitude motion, complex illumination change and low-texture areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video frame insertion, and in particular to a video frame insertion method based on an event-guided generative video diffusion model. Background Art

[0002] Video interpolation is an important research area in computer vision. Its main goal is to generate high-quality intermediate frames between existing video frames, thereby increasing the frame rate and temporal resolution of the video. High frame rate videos have a wide range of applications in filmmaking, virtual reality, slow-motion replay, and motion analysis.

[0003] Traditional video frame insertion methods are mainly divided into the following categories:

[0004] Optical flow-based methods estimate the pixel displacement field between adjacent frames and interpolate the intermediate moments. Representative methods include SuperSloMo (Jiang et al., CVPR 2018) and RIFE (Huang et al., CVPR 2022).

[0005] Direct generation methods based on deep learning: predict intermediate frames directly from input frames through end-to-end neural networks, such as CAIN (Choi et al., AAAI 2020) and XVFI (Sim et al., ICCV 2021).

[0006] Event camera-based methods: Utilize the high temporal resolution information captured by event cameras to assist in interpolation, such as TimeLens (Tulyakov et al., CVPR 2021) and CBMNet (Kim et al., CVPR 2023).

[0007] Diffusion model-based methods: Apply generative diffusion models to implement video interpolation, such as DualSVD (Wang et al., ICLR 2024).

[0008] In optical flow-based methods, optical flow estimation is often inaccurate in scenes with large motion, resulting in severe ghosting and blurring artifacts. They also lack the ability to model nonlinear motion, perform poorly in complex motion scenes, and are unable to handle occlusions and newly appearing objects, resulting in unnatural stretching and deformation.

[0009] In event camera-based methods, the feature extraction structure is simple, the high temporal resolution characteristics of event data are not fully utilized, the adaptability to large-scale motion and low-light conditions is insufficient, strong visual priors are lacking, the quality of reconstructed details is low, and the generalization ability is limited due to the small amount of existing event-RGB data.

[0010] Although the diffusion model-based method has good visual generation quality, it lacks precise timing guidance, the motion trajectory does not conform to physical laws, the bidirectional generation strategy has high computational cost and unstable training, and lacks a mechanism for combining with event data, and cannot utilize high temporal resolution information.

[0011] In summary, traditional video interpolation methods, limited by the inherent sampling frame rate of RGB cameras, often produce artifacts such as blurring, ghosting, or position errors in high-speed motion scenes. While existing event-assisted interpolation methods can obtain high-temporal resolution information, they still face problems such as insufficient model generalization and limited ability to model complex scenes. They perform poorly when dealing with large-scale motion, complex lighting changes, and low-texture areas. Summary of the Invention

[0012] The purpose of the present invention is to provide a video interpolation method based on an event-guided generative video diffusion model in order to overcome the problem that existing video interpolation technology has poor performance when processing large-scale motion, complex lighting changes and low-texture areas.

[0013] The purpose of the present invention can be achieved by the following technical solutions:

[0014] A video frame insertion method based on an event-guided generative video diffusion model, the method comprising the following steps:

[0015] Constructing an event-guided generative video diffusion network, the event-guided generative video diffusion network includes a conditional generator MMCG, a pre-trained VAE encoder / decoder, a spatiotemporal attention conditional U-network Unet, and a guided embedding feature aggregation module GEA. The conditional generator MMCG includes a region of interest (ROI) selector, a voxel feature extractor (VFE), a multimodal feature fusion module (MMF), and a three-dimensional stereo residual learning module (3DVR).

[0016] A two-stage training strategy is adopted. The specific two-stage training strategy is:

[0017] Obtain a training video dataset and input it into the conditional generator MMCG for the first stage of training. During the first stage of training, the parameters of the pre-trained VAE encoder / decoder and the guided embedding feature aggregation module GEA remain unchanged. The parameters of the voxel feature extractor VFE, the multimodal feature fusion module MMF, the 3D stereo residual learning module 3DVR, and the parameters of the 3D convolution module on the pre-trained VAE encoder side are trained.

[0018] Then, the second stage of training is carried out. Gaussian noise is added to the training latent features and normalized. Then, the denoised latent features are input into the spatiotemporal attention conditional U-type network (Unet). The denoised prediction values are calculated, and the spatiotemporal attention conditional U-type network (Unet) is fine-tuned by constructing a loss function.

[0019] After the two-stage training strategy is completed, an event-guided generative video diffusion model is obtained. The event-guided generative video diffusion model includes the generator MMCG trained in the first stage, the spatiotemporal attention conditional U-type network Unet trained in the second stage, and the pre-trained VAE encoder / decoder and guided embedding feature aggregation module GEA. Video data and event stream data are input into the guided generative video diffusion model to obtain the interpolated video.

[0020] Furthermore, the specific steps of inputting the training video dataset into the condition generator MMCG and performing the first stage training are as follows:

[0021] The training video dataset includes event stream E and video frames. The event stream E is input into the region of interest (ROI) selector after voxel grid encoding. The output of the ROI selector is used as the input of the voxel feature extractor (VFE). The voxel feature extractor (VFE) outputs the event stream latent space features.

[0022] Two input frames in the video frame are respectively input into the pre-trained VAE encoder, and after the replication and 3D convolution modules, two video frame latent space features are obtained. The event stream latent space features and the video frame latent space features are equal in the time dimension;

[0023] The event stream latent space features and the video frame latent space features are input into the multimodal feature fusion module MMF, and the fusion feature F is obtained after fusion. mmf Input 3D residual learning module 3DVR, 3D residual learning module 3DVR generates enhanced output features based on weighted fusion of fusion features and VAE latent features, VAE latent features represent the features of pre-trained VAE encoder and copy, the enhanced output feature f(t) is the coding condition c generated by the guidance of the condition generator MMCG output t , which is the predicted feature of the conditional generator, and then calculate the loss function to optimize the parameters of the voxel feature extractor VFE, the multimodal feature fusion module MMF, the three-dimensional stereo residual learning module 3DVR and the parameters of the 3D convolution module. During the first stage of training, the parameters of the pre-trained VAE encoder / decoder and the guided embedding feature aggregation module GEA remain unchanged.

[0024] Furthermore, the voxel feature extractor VFE outputs the event stream latent space features as follows:

[0025] The voxel feature extractor (VFE) obtains the voxelized event data output by the region of interest (ROI) selector and uses a three-level downsampling structure to gradually extract the spatiotemporal features F of the event data. e As the event stream latent space feature;

[0026] The spatiotemporal feature F e for:

[0027] F e =P3(F3(P2(F2(P1(F1(x))))))

[0028] Where x is the voxelized event data, F i represents the feature extraction operation of the i-th downsampling block, P i represents a three-dimensional average pooling operation with a kernel size of (1,2,2). Each downsampling block consists of a 3D convolutional layer and a LeakyReLU activation function.

[0029] Furthermore, the fusion feature F output by the multimodal feature fusion module MMF is mmf for:

[0030] F mmf =w1·F1+w2·F e +w3·F2+F fuse

[0031] Among them, F fuse is the fusion feature generated by additional 3D convolution, w1, w2 and w3 are the feature weights generated by the two-layer convolution network and adaptive average pooling, where

[0032] w i =Softamx(Conv3D(σ(Conv3D(P global (F cat )))))

[0033] P global represents global average pooling, σ is the LeakyReLU activation function, Conv3D is 3D convolution, F cat are the video frame latent space features F1 and F2 and the event stream latent space feature F e splicing features.

[0034] Furthermore, the enhanced output features generated by the 3D residual learning module 3DVR based on the weighted fusion of fusion features and VAE potential features are:

[0035] f(t)=w evs (t)·f evs +w1(t)·f1+w2(t)·f2

[0036] Among them, wevs (t), w1(t) and w2(t) are dynamic weights based on time series, f evs Based on the fusion human feature F mmf The final features generated, f1 and f2 are VAE latent features, which are the features obtained by copying two input frames into the pre-trained VAE encoder respectively;

[0037] The final features are:

[0038] f evs =Conv3D(x res +F mnf )

[0039] where x res Indicates F mmf Residual features learned by 5 cascaded 3D residual blocks.

[0040] Furthermore, the dynamic weight based on time series is:

[0041] w1(t)=t / T

[0042] w2(t)=(Tt) / T

[0043] w evs (t)=1,if t∈{1,2,…,T-1};w evs (t)=0,if t∈{0,T}

[0044] T represents the multiple of frame rate expansion. When a 30 frame rate video is expanded to 240 frames, T=8.

[0045] Furthermore, Unet predicts the potential features after denoising as:

[0046]

[0047] in, is the standardized training latent feature, c t Represents the encoding condition output by the condition generator MMCG;

[0048] The denoised prediction value is:

[0049]

[0050] Among them, z pred Predict the denoised potential features for Unet, represents the potential features for training after adding Gaussian noise, σ t is the noise level;

[0051] The loss function during fine-tuning is:

[0052]

[0053] Among them E z,N Indicates the expected value or average value.

[0054] Furthermore, when fine-tuning the spatiotemporal attention conditional U-network Unet, Unet includes a temporal self-attention layer, in which the spatiotemporal tensor X is reshaped into a new tensor X ′ , where B is the batch size, N is the number of frames, H and W are the spatial sizes, and C is the number of channels. The features of the temporal self-attention layer query Q, key K, and value V are:

[0055] Q=W q (X ′ )

[0056] K=W k (X ′ )

[0057] V=W v (X ′ )

[0058] Where W q 、W k and W v is the corresponding weight. When fine-tuning, only the time domain related W is adjusted v and W o The weight parameters of the Linear layer.

[0059] Furthermore, the guided embedding feature aggregation GEA includes a pre-trained CLIP model, the guided embedding feature aggregation GEA obtains an input frame, uses the pre-trained CLIP model to extract features of the input frame, and performs weighted average fusion on the features of the input frame to obtain a semantic feature e fused , semantic feature e fused Embedded into Unet as part of Unet.

[0060] Furthermore, the specific steps of inputting the video data and event stream data into the guided generative video diffusion model to obtain the inserted video are as follows:

[0061] Video data and event stream data are input into the conditional generator MMCG in the guided generative video diffusion model. Initialized Gaussian noise is added to the output of the conditional generator MMCG, and then input into the fine-tuned Unet for gradual denoising through DDIM iteration:

[0062] z t-1 =J(z t ,f θ (z t ;t,c t);t)

[0063] where f θ is the trained noise prediction network, J represents the denoising conversion operator, z t-1 is the feature after denoising in step t-1, c t represents the output of the condition generator MMCG, z t represents the initialized Gaussian noise;

[0064] After DDIM iteration and gradual denoising, the final potential representation z0 is obtained, and the final potential representation z0 is input into the pre-trained VAE decoder to obtain the interpolated video.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] This paper effectively combines the precise timing information of event streams with the powerful generation capabilities of diffusion models through an innovative multimodal condition generator and a selective temporal self-attention fine-tuning strategy. EGVD's core innovations include a multimodal motion condition generator (MMCG), a selective temporal self-attention fine-tuning strategy, and an EDM-based diffusion training framework. Through a two-stage training end-to-end framework, high-quality event-guided video interpolation is achieved. The present invention overcomes the limitations of existing video interpolation technology by combining a multimodal motion condition generator (MMCG) with a selective temporal self-attention fine-tuning strategy. The four core submodules of the MMCG (ROI selector, VFE, MMF, and 3DVR) effectively integrate the high temporal resolution information of the event camera with RGB frame data, enabling the system to accurately capture motion trajectories in large-scale motion scenes; selective temporal self-attention fine-tuning retains the powerful visual priors of the diffusion model while enhancing temporal relationship modeling; the EDM training framework ensures the stability of the model under different noise levels and improves the generation quality of low-texture areas; the two-stage training strategy reduces the overall training difficulty and ensures that the generated intermediate frames not only meet the precise timing constraints provided by the event data but also have high visual quality, thereby comprehensively improving the interpolation performance in challenging scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 Schematic diagram of the generative video frame insertion framework structure of the present invention;

[0068] Figure 2 This is the structure diagram of the voxel feature extraction module;

[0069] Figure 3 This is the structure diagram of the multimodal feature fusion module;

[0070] Figure 4 This is the structural diagram of the 3D residual learning module;

[0071] Figure 5The architecture diagram of the guided embedding feature aggregation module;

[0072] Figure 6 Comparison of visual quality between EGVD and other methods;

[0073] Figure 7 Compare the quality of different interpolation methods in multiple real scenes. DETAILED DESCRIPTION

[0074] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0075] This invention aims to address the poor performance of existing video interpolation techniques when dealing with large motion, complex illumination variations, and low-texture areas. Specifically, traditional video interpolation methods, limited by the inherent sampling frame rate of RGB cameras, often produce artifacts such as blurring, ghosting, and misaligned positions in high-speed motion scenes. Existing event-assisted interpolation methods, while capable of acquiring high-temporal resolution information, still suffer from insufficient model generalization and limited ability to model complex scenes. By combining the high temporal resolution of event cameras with the powerful priors of generative diffusion models, this invention provides a high-quality video interpolation solution that is both physically realistic and visually natural. This invention proposes an event-guided generative video diffusion model (EGVD). This model effectively combines the precise temporal information of event streams with the powerful generative capabilities of diffusion models through an innovative multimodal conditional generator and a selective temporal self-attention fine-tuning strategy. The core innovations of EGVD include a multimodal motion conditional generator (MMCG), a selective temporal self-attention fine-tuning strategy, and an EDM-based diffusion training framework. High-quality event-guided video interpolation is achieved through an end-to-end framework with two-stage training.

[0076] The event-guided generative video diffusion model (EGVD) proposed in this invention is mainly composed of a multimodal motion conditional generator (MMCG), a pre-trained VAE encoder / decoder, a spatiotemporal attention conditional U-network (Unet), and a guided embedding feature aggregation module. EGVD receives two frames of video images (I_t, I_{t+1}) and the event stream data E_{t→t+1} between them. The event stream is processed by MMCG to generate a multimodal motion conditional signal. The conditional signal is input into the spatiotemporal attention conditional Unet together with the potential features encoded by VAE. The video frame potential representation is generated through the denoising process, and finally a high-quality interpolated frame is obtained through the VAE decoder. The event-guided generative video interpolation framework structure diagram is shown in the figure below. Figure 1The figure shows the overall architecture of EGVD, including the multimodal motion condition generator, VAE encoder / decoder, spatiotemporal attention condition Unet, and inference process. The figure shows how to receive RGB frames and event streams and generate interpolated frames through full model processing. The figure shows the complete process of the present invention:

[0077] The framework first receives two frames of video images and the event stream data between them. The event stream first passes through the voxel grid encoding module in the data importer, which divides the event stream into a corresponding number of segments in chronological order. The number of these segments is determined by the required number of interpolation frames. The voxel grid of each segment is further divided into multiple channels according to time to better preserve the high temporal resolution characteristics of the event. In this invention, 8 channels are used to more accurately capture motion details. Next, the event stream passes through the region of interest (ROI) selector, which calculates the motion area in the current frame and selects the area related to the interpolation task, thereby enhancing the model's perception of the motion area. Subsequently, the event stream passes through the voxel feature extractor (VFE) to encode the spatiotemporal information into latent space features. The VFE module compresses the spatiotemporal information in the event stream into a low-resolution latent space, providing a compact and efficient representation for subsequent interpolation tasks. The latent space features of the event stream and the input frame are consistent in dimension. Specifically, the input image frame is encoded into latent space features by a variational autoencoder (VAE) encoder and then replicated by the number of output frames, ensuring that the latent space features of the event stream align with those of the input frame, enabling better feature fusion and temporal modeling. Within the SVD framework, the VAE is used to encode the input image and event stream into latent space features, providing a compact feature representation for subsequent interpolated frame generation. By importing the VAE's pre-trained weights, its powerful capabilities in feature extraction and latent space representation are fully utilized. The multimodal feature fusion (MMF) module effectively fuses the event stream features with those of the input image, providing rich temporal information for the subsequent 3D stereo residual learning (3DVR) module. The 3DVR module further optimizes the feature representation and enhances spatiotemporal consistency by reconstructing the encoding conditions that guide the generation. Building on these modules, the guided embedding feature aggregation (GEA) module further aggregates the features to generate the final embedding features. The parameters of the GEA module maintain the original structure of the SVD model and do not participate in any parameter updates during training. Finally, the spatiotemporal attention conditional UNet model is fine-tuned to further optimize the interpolated frame results. During fine-tuning, the model utilizes a forward diffusion process and a log-Gaussian sampling strategy to accelerate training. This process increases the probability of sampling low-level noise, enabling the model to train effectively in a low-noise environment. This accelerates fine-tuning convergence and improves detail generation. After the first phase of training and the second phase of fine-tuning are complete, interpolation inference can be performed. The input data is converted into embedding and encoding conditions. The fine-tuned spatiotemporal attention-conditioned Unet model is then used to perform the inverse denoising process within the latent space. After inverse denoising, the generated features are decoded into the final video frame using a VAE decoder, completing the interpolation task.

[0078] Multimodal Motion Condition Generator (MMCG), MMCG is the core innovative module of this invention, responsible for integrating event information and RGB frame information to generate conditional representations for controlling the diffusion process. MMCG consists of four main submodules:

[0079] Region of Interest (ROI) Selector: The ROI selector is designed to accurately extract the motion region related to the interpolation task from the event stream. The specific process is as follows:

[0080] 1. Normalize the input event stream E:

[0081] 2. Use Gaussian kernel G σ Smoothing the event image: E″=G σ *E′

[0082] 3. Convert the smoothed image into a binary image B through thresholding: B = 1, if E > 0.01; 0, otherwise

[0083] 4. Apply morphological dilation operation and median filtering to optimize the target area:

[0084] For each time step in the interpolation task, the ROI masks of multiple time channels are combined to obtain a unified motion region mask. This precise region selection mechanism significantly enhances the model's ability to perceive motion regions.

[0085] Voxel Feature Extractor (VFE), the structure of the voxel feature extraction module is as follows Figure 2 The figure shows the three-level downsampling structure of the VFE, including 3D convolutional layers, activation functions, and pooling operations. The figure clearly labels the number of channels, kernel size, and feature transformation process of each layer. The VFE module is responsible for extracting spatiotemporal features from voxelized event data. The VFE module uses a three-level downsampling structure to gradually extract the spatiotemporal features of the event data:

[0086] F e =P3(F3(P2(F2(P1(F1(x))))))

[0087] Among them F i represents the feature extraction operation of the i-th downsampling block, P i Represents a 3D average pooling operation with a kernel size of (1,2,2). Each downsampling block consists of a 3D convolutional layer and a LeakyReLU activation function.

[0088] In its implementation, VFE converts 8-channel voxel event data into 128-channel features after three levels of downsampling, reducing the spatial size to 1 / 8 of the original. This design effectively compresses the volume of event data while preserving key spatiotemporal dynamic information.

[0089] Multimodal feature fusion module (MMF), the structure of the multimodal feature fusion module is as follows Figure 3 The figure shows the adaptive two-way fusion structure of MMF, including feature concatenation, weight generation, and weighted fusion. The figure clearly shows how features from different modalities are fused through adaptive weights. The design goal of the MMF module is to effectively integrate features from event streams and RGB images. The MMF implementation is divided into the following steps:

[0090] 1. Feature splicing: Combine the two input frame features F1, F2 and the event stream feature F e Concatenate in the channel dimension

[0091] F2=Conv3D(R(E vAE (I0)))

[0092] F1=Conv3D(R(E vAE (I1)))

[0093] F e =E VFE (E 0→1 )

[0094] F cat =[F1,F e ,F2]

[0095] I0 and I1 represent the input images, E 0→1 is the input event stream, and R represents replication in the time dimension so that the image features encoded by VAE and the event features encoded by VFE are equal in the time dimension.

[0096] 2. Adaptive weight learning: Generate feature weights through a two-layer convolutional network and adaptive average pooling

[0097] w=Softmax(Conv3D(σ(Conv3D(P global (F cat )))))

[0098] Among them, P global represents global average pooling, and σ is the LeakyReLU activation function.

[0099] 3. Weighted feature fusion:

[0100] F mmf =w1·F1+w2·F e+w3·F2+F fuse

[0101] Among them F fuse It is a fusion feature generated by additional 3D convolution.

[0102] Through this adaptive two-way fusion strategy, the MMF module can dynamically adjust the importance of different features according to the characteristics of the input content to achieve the optimal feature fusion effect.

[0103] 3D residual learning module (3DVR), the structure of the 3D residual learning module is as follows Figure 4 As shown,

[0104] This figure illustrates the 3DVR architecture, including the residual block, temporal attention, spatial attention, and dynamic weight allocation mechanism. The figure shows the connections between these components and the direction of feature flow. The 3DVR module optimizes the fusion expression of multimodal features through residual learning. The 3DVR module includes core components: a cascade of five three-dimensional residual blocks (ResBlock3D), a temporal attention mechanism (TA), and a spatial attention mechanism (SA).

[0105] The processing flow is as follows:

[0106] 1. Features are enhanced layer by layer through cascaded ResBlock3D

[0107] 2. TA and SA are introduced after the three middle residual blocks to enhance spatiotemporal context perception:

[0108] f t =TA(x)·x res3d

[0109] f s =SA(f t )·f t

[0110] x res3d =f s

[0111] x res3d It is the residual feature output by the previous layer ResBlock3D

[0112] 3. Generate the final features through residual connection and convolution operation:

[0113] f evs =Conv3D(x res +F mmf )

[0114] xres It is the final residual feature of the 5 cascaded ResBlock3D outputs.

[0115] In addition, the 3DVR module also designs a dynamic weight allocation mechanism based on time sequence, which calculates the feature weight according to the relative time position t∈[0,T]:

[0116] w1(t)=t / T

[0117] w2(t)=(Tt) / T

[0118] w evs (t)=1,if t∈{1,2,…,T-1};w evs (t)=0,if t∈{0,T}

[0119] T represents the multiple of frame rate expansion. For example, when a 30-frame video is expanded to 240 frames, T=8.

[0120] Finally, the enhanced output features are generated through weighted fusion:

[0121] f(t)=w evs (t)·f evs +w1(t)·f1+w2(t)·f2

[0122] This design ensures a smooth transition in the time dimension while fully utilizing the contribution of event features to intermediate frames.

[0123] f1=R(E VAE (I1))

[0124] f2=R(E VAE (I0))

[0125] Guided Embedding Feature Aggregation (GEA), the architecture of the guided embedding feature aggregation module is as follows Figure 5 The figure shows the structure of GEA, including the CLIP encoder and the feature fusion process. The figure shows how semantic features are extracted from the input frames and fused. The GEA module uses the semantic understanding capabilities of the pre-trained CLIP model to provide semantic guidance for subsequent video frame generation. The specific implementation is as follows:

[0126] 1. Extract features from the input frame using the CLIP visual encoder:

[0127] e0=E CLIP (I0)

[0128] e N+1 =e CLIP (I N+1 )

[0129] 2. Perform a simple weighted average fusion of the two feature maps:

[0130] e fused =0.5(e0+e N+1 )

[0131] This design maintains the semantic consistency of features, avoids the distortion of semantic information that may be caused by complex fusion operations, and effectively guides the generation model to focus on the key semantic areas in the image. fused Embedded into Unet as part of the spatiotemporal attention conditional Unet, it provides semantic information to the generation process of the model.

[0132] Temporal self-attention fine-tuning,To efficiently adapt to the video interpolation task, this paper proposes a selective temporal self-attention fine-tuning strategy.,The core mechanism of the temporal self-attention layer is as follows:

[0133] 1. Transform the space-time tensor X∈R B×N×H×W×C Reshape into X ′ ∈R B×HW×N×C , where B is the batch size, N is the number of frames, H and W are the spatial dimensions, and C is the number of channels.

[0134] 2. Calculate the features of query (Q), key (K), and value (V):

[0135] Q=W q (X ′ )

[0136] K=W k (X ′ )

[0137] V=W v (X ′ )

[0138] 3. Use scaled dot product attention to calculate attention output:

[0139]

[0140] 4. Through the output linear layer W o Processing the final result

[0141] During fine-tuning, we only adjust the time-domain related linear layers (W v and W o ), retaining the pre-trained knowledge of the SVD model in spatial feature extraction. This strategy significantly reduces the training data requirements and computational cost while improving the convergence speed.

[0142] The present invention adopts a two-stage training strategy, the first stage: condition generator training.

[0143] First, train the MMCG module, and the optimization target is the mean square error in the latent space:

[0144]

[0145] Among them E VAE is a VAE encoder with fixed parameters, I i is the i-th frame in the sequence, G Θ is a conditional generator with parameter Θ, G Θ (E,I0,I N+1 ) i represents the predicted features of the conditional generator for the i-th frame.

[0146] While keeping the parameters of the VAE encoder and GEA module fixed, the VFE, MMF, 3DVR and related convolutional layer parameters are optimized. This latent space training strategy reduces computational complexity while fully utilizing the semantic features of the pre-trained VAE encoder.

[0147] Phase 2: Diffusion model fine-tuning. The second phase fine-tunes the diffusion model based on the EDM framework. The specific process is as follows:

[0148] 1. Add Gaussian noise to the latent features, with a noise level σ t Sample from a lognormal distribution:

[0149]

[0150] σ t =exp(ε t )

[0151] ε t ~N(0.7,1.6)

[0152] 2. Normalize the latent features after adding noise:

[0153]

[0154] 3. Predict the latent features after denoising:

[0155]

[0156] 4. Calculate the denoised prediction value:

[0157]

[0158] 5. Use weighted mean square error as the loss function:

[0159]

[0160] This EDM-based training strategy ensures the stability of training under different noise levels while improving the generation quality.

[0161] In the inference process (Unet), EGVD uses the DDIM sampler for efficient denoising. The specific process is as follows:

[0162] 1. Input two frames of video images (I t ,I t+1 ) and event stream data E t→t+1

[0163] 2. Generate conditional signal c through MMCG t

[0164] 3. Initialize z from Gaussian noise t

[0165] 4. Denoising by DDIM iteration:

[0166] z t-1 =J(z t ,f θ (z t ;t,c t );t)

[0167] where f θ is the trained noise prediction network, J represents the denoising conversion operator

[0168] 5. Obtain the final potential representation z0 and get the interpolated frame through the VAE decoder:

[0169] I interp =E VAE -1 (z0)

[0170] In practical applications, the number of sampling steps is set to 50 to strike a balance between generation quality and computational efficiency.

[0171] The key protection points of the present invention include:

[0172] The overall framework of the event-guided generative video diffusion model (EGVD), including a multimodal motion-conditioned generator and a selective temporal self-attention fine-tuning strategy;

[0173] Multimodal Motion Condition Generator (MMCG), which includes a region of interest selector, a voxel feature extractor, a multimodal feature fusion module, and a 3D stereo residual learning module;

[0174] Region of interest (ROI) selection method based on normalization, Gaussian smoothing, thresholding and morphological operations;

[0175] Voxel Feature Extractor (VFE) with a three-stage downsampling structure;

[0176] Multimodal feature fusion module (MMF) based on adaptive weight learning mechanism;

[0177] A 3D stereo residual learning module (3DVR) that combines temporal and spatial attention mechanisms;

[0178] Dynamic weight allocation mechanism based on relative time position;

[0179] A selective temporal self-attention fine-tuning strategy that only adjusts the temporally correlated linear layers of the diffusion model;

[0180] A diffusion training framework based on the EDM framework, using normalized input and output and dynamic loss weights;

[0181] A two-stage training strategy consisting of conditional generator training and diffusion model fine-tuning.

[0182] The visual quality comparison of EGVD of the present invention and other methods is shown in the figure below: Figure 6 The figure shows a comparison of the interpolation results of EGVD and other methods in challenging scenes. The figure includes the comparison results of typical challenging scenes such as large-scale motion, low light, and low texture, which intuitively demonstrates the advantages of the present invention.

[0183] Compared with CBMNet, the proposed method has significantly improved generation quality (LPIPS improved by 43.4%), enhanced low-light adaptability (PSNR increased by 0.73dB), improved large-scale motion processing capability (PSNR of DJI dataset reaches 22.46dB vs. 16.24dB), and a more efficient feature extraction mechanism.

[0184] Compared with DualSVD, the present invention has the following advantages: improved physical consistency (motion trajectory is more consistent with physical laws), improved computational efficiency (computational cost is reduced by about 50%), enhanced training stability, and obvious perceptual quality advantages (LPIPS is improved by 31.4%).

[0185] The present invention has general advantages: strong adaptability to multiple scenarios (excellent performance in text areas, low-texture areas and low-light environments), strong generalization ability, excellent detail retention ability, and end-to-end design that facilitates practical application.

[0186] The quality comparison of different interpolation methods in multiple real scenes is as follows Figure 7 As shown, #i represents the frame index. The three examples shown from top to bottom are from the real datasets HQEVFI, BSRGB, and ERDS, respectively. It is important to note that all results were generated using uniform inference weights and no dataset-specific training was performed.

[0187] This paper experimentally validates the proposed method on four diverse datasets: Prophesee, BS-ERGB, DJI 30fps, and GOPRO 240fps. The main results are as follows:

[0188] Prophesee dataset: PSNR 22.87dB, SSIM 0.7768 (best), LPIPS 0.1422 (best, 27.4% improvement over the second-best method);

[0189] BS-ERGB dataset: PSNR 19.81dB, LPIPS 0.1732 (best, 24.1% improvement over the second-best method);

[0190] DJI 30fps dataset: PSNR 22.46dB (best, 3.73dB improvement over the second-best method), SSIM 0.7555 (best), LPIPS 0.1993 (best, 31.4% improvement over the second-best method);

[0191] GOPRO 240fps dataset: PSNR 23.84dB (best, 1.94dB higher than the second-best method), SSIM 0.8039 (best);

[0192] EGVD performs exceptionally well in challenging scenarios, such as text regions (PSNR improved by 7.09dB), low-texture regions (LPIPS improved by 58.8%), and low-light environments (LPIPS improved by 15.9%). Ablation experiments further validate the contribution of each module to overall performance. Removing the SVD denoiser, MMCG, VFE, MMF, 3DVR, or AW all results in varying degrees of performance degradation.

[0193] This invention can also be applied to a variety of fields, including video super-resolution, video deblurring, video restoration, slow-motion generation, 3D reconstruction, motion capture, and autonomous driving perception. In the field of video super-resolution, the high temporal resolution of event streams can be exploited to improve spatial resolution; in the field of video restoration, the diffusion model generation capability and event timing information can be used to repair damaged or missing video frames; and in the field of autonomous driving perception, the system's ability to identify and track fast-moving targets can be enhanced.

[0194] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A video frame insertion method based on an event-guided generative video diffusion model, characterized in that: The method comprises the following steps: Constructing an event-guided generative video diffusion network, the event-guided generative video diffusion network includes a conditional generator MMCG, a pre-trained VAE encoder / decoder, a spatiotemporal attention conditional U-network Unet, and a guided embedding feature aggregation module GEA. The conditional generator MMCG includes a region of interest (ROI) selector, a voxel feature extractor (VFE), a multimodal feature fusion module (MMF), and a three-dimensional stereo residual learning module (3DVR). A two-stage training strategy is adopted. The specific two-stage training strategy is: Obtain a training video dataset and input it into the conditional generator MMCG for the first stage of training. During the first stage of training, the parameters of the pre-trained VAE encoder / decoder and the guided embedding feature aggregation module GEA remain unchanged. The parameters of the voxel feature extractor VFE, the multimodal feature fusion module MMF, the 3D stereo residual learning module 3DVR, and the parameters of the 3D convolution module on the pre-trained VAE encoder side are trained. Then, the second stage of training is carried out. Gaussian noise is added to the training latent features and normalized. Then, the denoised latent features are input into the spatiotemporal attention conditional U-type network (Unet). The denoised prediction values are calculated, and the spatiotemporal attention conditional U-type network (Unet) is fine-tuned by constructing a loss function. After the two-stage training strategy is completed, an event-guided generative video diffusion model is obtained. The event-guided generative video diffusion model includes the generator MMCG trained in the first stage, the spatiotemporal attention conditional U-type network Unet trained in the second stage, and the pre-trained VAE encoder / decoder and guided embedding feature aggregation module GEA. Video data and event stream data are input into the guided generative video diffusion model to obtain the interpolated video.

2. The video frame insertion method based on the event-guided generative video diffusion model according to claim 1, characterized in that: The specific steps of inputting the training video dataset into the condition generator MMCG and performing the first stage training are as follows: The training video dataset includes event stream E and video frames. The event stream E is input into the region of interest (ROI) selector after voxel grid encoding. The output of the ROI selector is used as the input of the voxel feature extractor (VFE). The voxel feature extractor (VFE) outputs the event stream latent space features. Two input frames in the video frame are respectively input into the pre-trained VAE encoder, and after the replication and 3D convolution modules, two video frame latent space features are obtained. The event stream latent space features and the video frame latent space features are equal in the time dimension; The event stream latent space features and the video frame latent space features are input into the multimodal feature fusion module MMF, and the fusion feature F is obtained after fusion. mmf Input 3D residual learning module 3DVR, 3D residual learning module 3DVR generates enhanced output features based on weighted fusion of fusion features and VAE latent features, VAE latent features represent the features of pre-trained VAE encoder and copy, the enhanced output feature f(t) is the coding condition c generated by the guidance of the condition generator MMCG output t , which is the predicted feature of the conditional generator, and then calculate the loss function to optimize the parameters of the voxel feature extractor VFE, the multimodal feature fusion module MMF, the three-dimensional stereo residual learning module 3DVR and the parameters of the 3D convolution module. During the first stage of training, the parameters of the pre-trained VAE encoder / decoder and the guided embedding feature aggregation module GEA remain unchanged.

3. The video frame insertion method based on the event-guided generative video diffusion model according to claim 2, characterized in that: The voxel feature extractor VFE outputs the event stream latent space features as follows: The voxel feature extractor (VFE) obtains the voxelized event data output by the region of interest (ROI) selector and uses a three-level downsampling structure to gradually extract the spatiotemporal features F of the event data. e As the event stream latent space feature; The spatiotemporal feature F e for: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> e <h2 style=";text-align:left;direction:ltr"> =P3(F3(P2(F2(P1(F1(x)))))) Where x is the voxelized event data, F i represents the feature extraction operation of the i-th downsampling block, P i represents a three-dimensional average pooling operation with a kernel size of (1,2,2). Each downsampling block consists of a 3D convolutional layer and a LeakyReLU activation function.

4. The video frame insertion method based on the event-guided generative video diffusion model according to claim 3, characterized in that: The fusion feature F output by the multimodal feature fusion module MMF mmf for: F mmf =w1·F1+w2·F e +w3·F2+F fuse Among them, F fuse is the fusion feature generated by additional 3D convolution, w1, w2 and w3 are the feature weights generated by the two-layer convolution network and adaptive average pooling, where w i =Softmax(Conv3D(σ(Conv3D(P global (F cat ))))) P global represents global average pooling, σ is the LeakyReLU activation function, Conv3D is 3D convolution, F cat are the video frame latent space features F1 and F2 and the event stream latent space feature F e splicing features.

5. The video frame insertion method based on the event-guided generative video diffusion model according to claim 4, characterized in that: The enhanced output features generated by the 3D residual learning module 3DVR based on the weighted fusion of fusion features and VAE potential features are: f(t)=w evs (t)·f evs +w1(t)·f1+w2(t)·f2 Among them, w evs (t), w1(t) and w2(t) are dynamic weights based on time series, f evs Based on the fusion human feature F mmf The final features generated, f1 and f2 are VAE latent features, which are the features obtained by copying two input frames into the pre-trained VAE encoder respectively; The final features are: f evs =Conv3D(x res +F mnf ) where x res Indicates F mmf Residual features learned by 5 cascaded 3D residual blocks.

6. The video frame insertion method based on the event-guided generative video diffusion model according to claim 5, characterized in that: The dynamic weight based on timing is: w1(t)=t / T w2(t)=(Tt) / T w evs (t)=1,if t∈{1,2,…,T-1};w evs (t)=0,if t∈{0,T} T represents the multiple of frame rate expansion. When a 30 frame rate video is expanded to 240 frames, T=8.

7. The video frame insertion method based on the event-guided generative video diffusion model according to claim 1, characterized in that: Unet predicts the potential features after denoising as: in, is the standardized training latent feature, c t Represents the encoding condition output by the condition generator MMCG; The denoised prediction value is: Among them, z pred Predict the denoised potential features for Unet, represents the potential features for training after adding Gaussian noise, σ t is the noise level; The loss function during fine-tuning is: Among them E z,N Indicates the expected value or average value.

8. The video frame insertion method based on the event-guided generative video diffusion model according to claim 1, characterized in that: When fine-tuning the spatiotemporal attention conditional U-network Unet, Unet includes a temporal self-attention layer, in which the spatiotemporal tensor X is reshaped into a new tensor X ′ , where B is the batch size, N is the number of frames, H and W are the spatial sizes, and C is the number of channels. The features of the temporal self-attention layer query Q, key K, and value V are: Q=W q (X ′ ) K=W k (X ′ ) V=W v (X ′ ) Where W q 、W k and W v is the corresponding weight. When fine-tuning, only the time domain related W is adjusted v and W o The weight parameters of the linear layer Linear.

9. The video frame insertion method based on event-guided generative video diffusion model according to claim 1, characterized in that: The guided embedding feature aggregation GEA includes a pre-trained CLIP model. The guided embedding feature aggregation GEA obtains an input frame, extracts features of the input frame using the pre-trained CLIP model, and performs weighted average fusion on the features of the input frame to obtain a semantic feature e fused , semantic feature e fused Embedded into Unet as part of Unet.

10. The video frame insertion method based on event-guided generative video diffusion model according to claim 1, characterized in that: The specific steps of inputting video data and event stream data into the guided generative video diffusion model to obtain the interpolated video are as follows: Video data and event stream data are input into the conditional generator MMCG in the guided generative video diffusion model. Initialized Gaussian noise is added to the output of the conditional generator MMCG, and then input into the fine-tuned Unet for gradual denoising through DDIM iteration: z t-1 =J(z t ,f θ (z t ;t,c t );t) where f θ is the trained noise prediction network, J represents the denoising conversion operator, z t-1 is the feature after denoising in step t-1, c t represents the output of the condition generator MMCG, z t represents the initialized Gaussian noise; After DDIM iteration and gradual denoising, the final potential representation z0 is obtained, and the final potential representation z0 is input into the pre-trained VAE decoder to obtain the interpolated video.

Citation Information

Cited By

  • Long video editing method and system based on bidirectional attention and multi-mode guidance

    CN121334452A