Mono Video to Stereo Conversion via Depth-Aware Pixel Shifting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for generating stereo videos from mono videos are not robust, leading to incorrect object segmentation and inconsistent depth illusion, making it difficult to accurately depict the depth of objects in a frame.
Innovation Solution
A system that partitions mono videos into shots, determines depth parameters and pixel depth maps based on YCbCr color components, and shifts pixels in the left frames to create corresponding right frames, with optional scene classification and interpolation to improve quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If motion-based techniques are used to generate stereo video from mono video, then the process is simpler and less expensive, but object segmentation becomes incorrect and depth representation becomes inconsistent
Solution Approach 1:
The patent changes the parameter basis from motion information to color information (YCbCr color space). By using the Cb and Cr components which encode depth-related color variations, the system achieves accurate object segmentation and consistent depth representation without relying on motion analysis, thus resolving the contradiction between ease of generation and depth accuracy
Solution Approach 2:
The patent replaces the motion-based mechanical analysis system with a color-based optical analysis system. Instead of tracking object movement between frames to infer depth, the system directly extracts depth information from color components in the YCbCr color space, achieving more accurate and consistent depth representation
2Reliability
If pixel shifting is applied to create depth effect, then stereo video is generated, but empty pixels and edge artifacts are introduced
Solution Approach 1:
The patent applies preliminary action by performing interpolation to fill empty pixels before final stereo video output. The system predicts and fills in missing pixel values based on surrounding pixel data, preventing artifacts and ensuring complete, high-quality stereo video frames
Solution Approach 2:
The patent converts the harmful effect of empty pixels and edge artifacts into a benefit by using these regions as indicators for where interpolation is most needed. The system identifies areas affected by pixel shifting and applies targeted interpolation to smooth edges and fill gaps, transforming potential defects into opportunities for quality enhancement
Data Source
AI summary
A system and methodology provide for generation of a stereo video from a mono video. A mono video is partitioned into shots, where each shot including one or more frames of the mono video. The mono video frames are used as the left frames in the stereo video. Depth parameters are determined for each shot. A pixel depth map is created for each frame of each shot. The right frames for the stereo video are created by shifting pixels from the left frames laterally to occupy new locations in the right frames. A pixel's shift is based on the depth parameters and the pixel depth. In aggregate, pixel shifts will cause objects to appear in different locations in the right frame relative to where they appeared in the left frame. The effect upon the viewer is that the stereo video will provide an enhanced illusion of depth for the viewer.


