Separable Kernel Video Frame Interpolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video frame interpolation techniques face challenges such as high processing time, unwanted brightness changes, and lack of generalizability due to complex kernel-based methods and machine learning models that are not effectively trained to capture true motion between frames.
Innovation Solution
The proposed solution involves a kernel prediction network that estimates separable one-dimensional kernels for video frame interpolation, combined with kernel normalization and a contextual loss function to improve interpolation quality, and additional optimizations like delayed padding and self-ensembling to enhance processing efficiency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If kernel-based interpolation methods are used to improve interpolation quality, then processing time increases significantly
Solution Approach 1:
The patent segments the two-dimensional kernel into two separate one-dimensional kernels (horizontal and vertical). This segmentation reduces the computational complexity from O(n²) for a 2D kernel to O(n) for two 1D kernels, significantly decreasing processing time while maintaining interpolation quality through separable convolution operations.
Solution Approach 2:
The patent changes the parameter representation from full two-dimensional kernel coefficients to separable one-dimensional kernel coefficients. This parameter transformation reduces the number of parameters to be computed and stored, leading to faster processing while the normalization technique ensures the separable kernels maintain the desired interpolation properties.
2Adaptability or versatility
If machine learning models are trained to predict kernel coefficients, then model generalizability to arbitrary inputs is poor
Solution Approach 1:
The patent extracts the motion information directly from the input video frames themselves, rather than relying on pre-trained models to generalize from external training data. By using the input frames to directly guide the kernel prediction through optical flow estimation and separable convolution, the system achieves adaptability to arbitrary inputs without requiring extensive training time.
3Manufacturing precision
If larger kernels are used to improve interpolation quality, then memory demands increase quadratically
Solution Approach 1:
The patent segments the large two-dimensional kernel into two smaller one-dimensional kernels. This segmentation reduces memory requirements from storing an n×n matrix to storing two vectors of length n, decreasing memory demand from O(n²) to O(n) while preserving the ability to model large motions through the separable convolution process.
Data Source
AI summary
Embodiments are disclosed for video image interpolation. In some embodiments, video image interpolation includes receiving a pair of input images from a digital video, determining, using a neural network, a plurality of spatially varying kernels each corresponding to a pixel of an output image, convolving a first set of spatially varying kernels with a first input image from the pair of input images and a second set of spatially varying kernels with a second input image from the pair of input images to generate filtered images, and generating the output image by performing kernel normalization on the filtered images.


