Multi-View Interactive Media Guidance via IMU and Image Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality technologies face challenges in efficiently generating and implementing three-dimensional (3D) information for dynamic scenes, as the process of creating 3D reconstructions is computationally expensive and typically restricted to static environments.

Innovation Solution

The system utilizes inertial measurement units (IMUs) and image data to generate views of synthetic objects, allowing for the creation of immersive surround views by fusing 2D and 3D data, and providing interactive digital media representations that can model moving scenery and objects, while reducing computational costs through efficient data compression and stabilization techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If 3D reconstruction is used to add three-dimensional information to video and image data, then the quality and immersion of augmented reality is improved, but the computational cost and processing time increase significantly

Engineering Contradiction:
Improvequality of augmented realityVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system segments the scene into static background elements and dynamic foreground objects. 3D reconstruction is applied only to static elements using multi-view geometry, while dynamic objects are handled through video processing and motion estimation. This division reduces the overall computational burden while maintaining visual quality and immersion in the augmented reality experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from static 3D reconstruction to a dynamic approach where the scene is continuously updated based on camera motion and object movement. By using structure-from-motion algorithms and tracking dynamic objects across frames, the system adapts the level of 3D processing to the actual scene complexity, reducing unnecessary computational overhead while maintaining reliability.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If 3D reconstruction is performed on dynamic scenes, then the applicability of augmented reality to moving scenery is improved, but the computational complexity and processing requirements increase

Engineering Contradiction:
Improveapplicability to dynamic scenesVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs dynamic scene understanding through continuous tracking of moving objects and camera pose estimation. By distinguishing between static and dynamic elements in real-time and applying appropriate processing pipelines, the system achieves versatility in handling diverse scenes without requiring full 3D reconstruction of all elements, thus managing computational complexity effectively.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system replaces traditional mechanical 3D reconstruction pipelines with computational approaches based on multi-view geometry and structure-from-motion algorithms. These computational methods automatically infer 3D structure from 2D images without requiring explicit manual modeling, reducing device complexity while maintaining adaptability to dynamic scenes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If multiple camera views are captured to create immersive surround views, then the user experience is improved, but the data size and storage requirements increase

Engineering Contradiction:
Improveuser experience qualityVSAvoiddata size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system merges multiple camera views into a unified multi-view representation by aligning and integrating data from different angles. Through view synthesis and rendering techniques, the system combines redundant information across views while preserving unique visual content, achieving immersive surround views with optimized data size that balances user experience quality and storage requirements.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10665024B2Providing recording guidance in generating a multi-view interactive digital media representation
Publication Date: 2020.05.26 FUSION INC
  • US10665024B2 patent drawing
  • US10665024B2 patent drawing
  • US10665024B2 patent drawing

AI summary

Various embodiments of the present invention relate generally to systems and methods for collecting, analyzing, and manipulating images and video. According to particular embodiments, live images captured by a camera on a mobile device may be analyzed as the mobile device moves along a path. The live images may be compared with a target view. A visual indicator may be provided to guide the alteration of the positioning of the mobile device to more closely align with the target view.