Multi-View Interactive Media Conversion via Sensor-Guided Image Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional digital media formats, such as 2D flat images and videos, are limited in providing an immersive and interactive experience, failing to effectively capture and index visual data, especially with the increasing quantity of visual data being captured, which necessitates more comprehensive search and indexing mechanisms.
Innovation Solution
A system and method for converting a sequence of images captured by a mobile device into a multi-view interactive digital media representation (MVIDMR) that allows for 3D rotation of objects without a 3D polygon model, using sensor data from an inertial measurement unit to select images based on camera angles, and encoding them as a video for storage and playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional 2D flat images and videos are used, then the device complexity is low, but the immersion and interactivity are limited
Solution Approach 1:
The patent transitions from traditional 2D flat images to multi-view interactive representations that incorporate depth information and multiple camera angles, creating a 3D-like viewing experience without requiring full 3D polygon models. This dimensional enhancement provides immersion and interactivity while avoiding the complexity of complete 3D reconstruction
Solution Approach 2:
The system creates simplified copies of 3D scenes using 2D images from multiple cameras, generating multi-view representations that simulate 3D effects. These image-based models serve as lightweight alternatives to complex polygon models, providing interactive viewing experiences with significantly reduced computational requirements
2Productivity
If comprehensive search and indexing mechanisms are implemented, then the visual data management improves, but the system complexity increases
Solution Approach 1:
The patent segments visual data into structured multi-view representations with organized metadata including camera positions, timestamps, and spatial relationships. This segmentation enables efficient indexing and search operations by breaking down complex visual datasets into manageable, queryable components without requiring overly complex management systems
3Adaptability or versatility
If 3D polygon models are used to achieve 3D rotation, then the viewing experience is immersive, but the computational resources and manufacturing complexity increase
Solution Approach 1:
The system uses image-based models as simplified copies that enable 3D rotation effects without requiring full 3D polygon models. By synthesizing views from multiple 2D images captured at different angles, the system achieves rotational capability with significantly reduced computational resources and simpler processing requirements
Solution Approach 2:
The patent replaces the mechanical complexity of 3D polygon model construction and manipulation with a computational approach using 2D image synthesis. Instead of building and rotating complex 3D geometric models, the system synthesizes rotated views by processing and combining 2D images, substituting geometric computation with image processing operations that are computationally more efficient
Data Source
AI summary
Various embodiments of the present invention relate generally to systems and methods for analyzing and manipulating images and video. In particular, a multi-view interactive digital media representation can be generated from live images captured from a camera as the camera moves along a path. Then, a sequence of the images can be selected based upon sensor data from an inertial measurement unit and upon image data such that one of the live images is selected for each of a plurality of poses along the path. A multi-view interactive digital media representation may be created from the sequence of images, and the images may be encoded as a video via a designated encoding format.


