Multi-View Video Stabilization Using Depth Layer Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional digital video stabilization methods for multi-view systems, particularly those using rolling shutter type sensors, face challenges in distinguishing local motion from global motion and 3D motion from artifacts, leading to inefficient processing and potential image distortions, especially in real-time systems and when handling scenes with strong 3D nature.
Innovation Solution
The method involves acquiring simultaneous views from multiple proximate viewpoints, estimating global transformation parameters, and compensating for unintentional motion and global deformation using processor-implemented algorithms that selectively interpolate depth maps and differentiate between intentional and unintentional motion, allowing for real-time stabilization without the need for motion sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional digital video stabilization methods are used for multi-view systems with rolling shutter sensors, then image stabilization can be achieved, but computational cost increases significantly and real-time processing becomes difficult
Solution Approach 1:
The patent segments the video stabilization process into independent depth layers based on depth map information. By processing different depth layers separately and using the structure-from-motion results to guide the stabilization, the computational complexity is reduced while maintaining stabilization quality. This allows real-time processing by breaking down the complex global optimization into smaller, more manageable sub-problems.
Solution Approach 2:
The patent performs structure-from-motion analysis in advance to obtain 3D description of the scene and camera movement paths before performing the actual stabilization. This preliminary computation provides guidance for the subsequent stabilization process, reducing the computational burden during real-time video processing and enabling real-time operation.
2Measurement precision
If structure-from-motion and 3D camera movement determination are performed to improve stabilization accuracy, then stabilization precision improves, but computational cost becomes very high
Solution Approach 1:
The patent applies different processing quality and computational effort to different depth layers. By using depth information to identify which layers contain significant motion and which are static, the system concentrates computational resources on the most critical layers, reducing overall computational cost while maintaining high precision for the most important parts of the scene.
Solution Approach 2:
The patent uses depth maps as an intermediary to bridge the structure-from-motion analysis and the actual stabilization process. The depth information serves as a mediator that guides both processes, allowing the system to achieve accurate camera movement estimation without performing computationally expensive full 3D reconstruction and motion analysis for every pixel.
3Use of energy by moving object
If rolling shutter type sensors are used to increase sensitivity, then photon gathering capability improves, but distortions and artifacts occur during image acquisition
Solution Approach 1:
The patent replaces mechanical or optical correction methods with digital image processing. By using computational algorithms that analyze depth information and apply warping transformations in the digital domain, the system corrects rolling shutter distortions and artifacts without requiring complex optical mechanical systems, thus maintaining the sensitivity benefits of rolling shutter sensors while eliminating the distortion problems.
Data Source
AI summary
A system, method, and computer program product for digital stabilization of video data from cameras producing multiple simultaneous views, typically from rolling shutter type sensors, and without requiring a motion sensor. A first embodiment performs an estimation of the global transformation on a single view and uses this transformation for correcting other views. A second embodiment selects a distance at which a maximal number of scene points is located and considers only the motion vectors from these image areas for the global transformation. The global transformation estimate is improved by averaging images from several views and reducing stabilization when image conditions may cause incorrect stabilization. Intentional motion is identified confidently in multiple views. Local object distortion may be corrected using depth information. A third embodiment analyzes the depth of the scene and uses the depth information to perform stabilization for each of multiple depth layers separately.


