3D Image Visualization Using Camera Pose and Depth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current photo management applications lack the ability to generate and utilize 3D information for organizing, visualizing, and navigating digital images, resulting in a static and unengaging display experience.

Innovation Solution

The techniques involve generating 3D information by identifying and comparing visual word frequencies between images, using structure from motion (SfM) to create 3D representations, and employing 3D attributes like camera pose and position to organize and display images in a dynamic and interactive manner.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If 3D information is generated and used to organize and display images, then the user interaction and immersion are enhanced, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transitions from traditional 2D flat image layouts to 3D spatial arrangements by generating depth information from image pairs and using stereo visualization techniques. This dimensional change enables more immersive and interactive user experiences while organizing images in three-dimensional space based on captured location and orientation data.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system performs preliminary 3D reconstruction and depth map generation during the image processing stage, creating pre-computed 3D representations that can be efficiently rendered and navigated later. This preliminary action reduces real-time processing requirements during user interaction.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If visual word frequencies are compared to identify related images, then image grouping accuracy is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveimage matching accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the image analysis process into distinct stages: extracting visual words from images, computing frequency distributions, comparing histograms, and generating similarity scores. This segmentation allows for optimized processing at each stage and enables parallel computation of visual word frequencies across multiple images.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transforms image data into a different parameter space by converting visual features into visual word frequencies and histogram representations. This parameter transformation enables efficient comparison and matching by working with compressed frequency distributions rather than raw pixel data.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9311756B2Image group processing and visualization
Publication Date: 2016.04.12 APPLE INC
  • US9311756B2 patent drawing
  • US9311756B2 patent drawing
  • US9311756B2 patent drawing

AI summary

Techniques are provided for efficiently generating 3D information from a set of digital images. Techniques are also provided for displaying groups (or clusters) of digital images using 3D information associated with the digital images. In one technique, a group of digital images are displayed as a stack of thumbnail images where the thumbnail images are aligned on a display with respect to each other based on common features identified in the digital images, camera position, and/or camera pose. In another technique, a group of digital images are organized on a display in either a 3D layout or a 2D layout based on 3D information associated with each digital image in the group. In another technique, a transition effect is generated based on projections of two digital images onto a common scene plane and blending (or cross fading) one of the 3D projections with the other of the 3D projections.