Video Frame Interpolation Using Layered Motion Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Live streaming businesses face high bandwidth costs due to the need for high frame rate videos, and distributing low frame rate videos results in inferior viewing experiences compared to 60 FPS or 120 FPS videos.
Innovation Solution
A method and apparatus for video frame interpolation that acquires motion level information and deep frame interpolation features of adjacent frames to generate intermediate frames, using a layered perception technology that classifies pixels by motion levels and performs frame interpolation layer by layer, reducing the complexity of optical flow estimation and alleviating issues like image distortion and jitter.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If low frame rate videos are distributed to reduce bandwidth costs, then bandwidth costs are reduced, but viewing experience deteriorates
Solution Approach 1:
The patent performs frame interpolation in advance to generate intermediate frames between key frames, so that when low frame rate videos are distributed, the viewing experience is enhanced without requiring high frame rate source material. This preliminary processing allows standard definition videos to be played back smoothly at high frame rates.
Solution Approach 2:
The patent introduces intermediate frames as a mediator between key frames to bridge the temporal gap. These intermediate frames are generated through optical flow estimation and pixel-level interpolation, serving as a bridge that connects discrete key frames to create continuous high frame rate playback.
2Productivity
If traditional optical flow estimation is used for frame interpolation, then frame interpolation can be achieved, but image distortion and jitter occur
Solution Approach 1:
The patent segments the image into multiple layers based on motion levels, where each layer contains pixels with similar motion characteristics. This segmentation allows different interpolation strategies to be applied to different layers, preventing the mixing of static and dynamic elements that causes distortion and jitter in traditional methods.
Solution Approach 2:
The patent applies different interpolation quality levels to different regions of the image based on their motion characteristics. Static regions receive simpler interpolation while dynamic regions receive more sophisticated processing, optimizing both quality and computational efficiency while maintaining image stability.
3Device complexity
If layered perception technology is used to classify pixels by motion levels, then optical flow estimation complexity is reduced and image distortion is alleviated, but processing steps increase
Solution Approach 1:
The patent divides pixels into multiple motion levels (typically 3-5 levels) based on their motion characteristics. This segmentation simplifies optical flow estimation by allowing each level to be processed independently with appropriate algorithms, reducing overall computational complexity while maintaining accuracy.
Solution Approach 2:
The patent dynamically adjusts the number of motion levels and processing parameters based on the specific video content and motion characteristics. This allows the system to adapt to different scenarios, using fewer levels for simple content and more levels for complex motion, balancing complexity and effectiveness.
Data Source
AI summary
A method and apparatus for video frame interpolation, and a device and a storage medium for the same are provided. The method may include: acquiring a target video, and acquiring a (t−1)th frame of image and a tth frame of image in the target video; acquiring motion level information of pixel points of the (t−1)th frame of image and the tth frame of image; acquiring deep frame interpolation features of the (t−1)th frame of image and the tth frame of image, respectively; and performing a frame interpolation operation layer by layer based on the deep frame interpolation features and the motion level information of the (t−1)th frame of image and the tth frame of image to generate an intermediate frame between the (t−1)th frame of image and the tth frame of image, and interpolating the intermediate frame between the (t−1)th frame of image and the tth frame of image.


