Separable Kernel Video Frame Interpolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video frame interpolation techniques face challenges such as high processing time, unwanted brightness changes, and lack of generalizability due to complex kernel-based methods and machine learning models that are not effectively trained to capture true motion between frames.

Innovation Solution

The proposed solution involves a kernel prediction network that estimates separable one-dimensional kernels for video frame interpolation, combined with kernel normalization and a contextual loss function to improve interpolation quality, and additional optimizations like delayed padding and self-ensembling to enhance processing efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If kernel-based interpolation methods are used to improve interpolation quality, then processing time increases significantly

Engineering Contradiction:
Improveinterpolation qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the two-dimensional kernel into two separate one-dimensional kernels (horizontal and vertical). This segmentation reduces the computational complexity from O(n²) for a 2D kernel to O(n) for two 1D kernels, significantly decreasing processing time while maintaining interpolation quality through separable convolution operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation from full two-dimensional kernel coefficients to separable one-dimensional kernel coefficients. This parameter transformation reduces the number of parameters to be computed and stored, leading to faster processing while the normalization technique ensures the separable kernels maintain the desired interpolation properties.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If machine learning models are trained to predict kernel coefficients, then model generalizability to arbitrary inputs is poor

Engineering Contradiction:
Improvemodel generalizabilityVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent extracts the motion information directly from the input video frames themselves, rather than relying on pre-trained models to generalize from external training data. By using the input frames to directly guide the kernel prediction through optical flow estimation and separable convolution, the system achieves adaptability to arbitrary inputs without requiring extensive training time.

Inventive Principle:
Principle #2Taking out (Extraction)

3Manufacturing precision

If larger kernels are used to improve interpolation quality, then memory demands increase quadratically

Engineering Contradiction:
Improveinterpolation qualityVSAvoidmemory demand
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the large two-dimensional kernel into two smaller one-dimensional kernels. This segmentation reduces memory requirements from storing an n×n matrix to storing two vectors of length n, decreasing memory demand from O(n²) to O(n) while preserving the ability to model large motions through the separable convolution process.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11871145B2Optimization of adaptive convolutions for video frame interpolation
Publication Date: 2024.01.09 ADOBE INC
  • US11871145B2 patent drawing
  • US11871145B2 patent drawing
  • US11871145B2 patent drawing

AI summary

Embodiments are disclosed for video image interpolation. In some embodiments, video image interpolation includes receiving a pair of input images from a digital video, determining, using a neural network, a plurality of spatially varying kernels each corresponding to a pixel of an output image, convolving a first set of spatially varying kernels with a first input image from the pair of input images and a second set of spatially varying kernels with a second input image from the pair of input images to generate filtered images, and generating the output image by performing kernel normalization on the filtered images.