Multi-view Interactive Digital Media Representation for 3D Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional digital media formats, such as 2D flat images and videos, are limited in providing an immersive and interactive experience, especially when trying to capture and present three-dimensional data, as they require significant additional data for interpolation or extrapolation, leading to high processing demands and inefficient data transfer.

Innovation Solution

The development of multi-view interactive digital media representations (MVIDMR) that analyze spatial relationships between images and video with location information to create an immersive, interactive, and efficient format for capturing and presenting 3D data, allowing for the generation and tagging of media content associated with specific locations within the media.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional 2D flat images and videos are used to capture and present 3D data, then the format is simple and widely compatible, but significant additional data is required for interpolation or extrapolation, leading to high processing demands and inefficient data transfer

Engineering Contradiction:
Improve3D data representation capabilityVSAvoidprocessing demands
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transitions from 2D flat images to multi-view 3D representations by capturing images from multiple angles and depths. This dimensional expansion enables true 3D data representation while using efficient encoding methods that avoid the need for extensive interpolation or extrapolation, thereby reducing processing demands despite the increased adaptability for 3D content

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent divides the 3D scene into multiple discrete views captured from different positions and angles. Each view is processed and encoded independently, allowing for efficient data management and transfer. This segmentation approach reduces the overall processing complexity compared to attempting to represent the entire 3D scene as a single interpolated model

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If dense depth maps or optical flow maps are used to describe scene structure, then 3D reconstruction accuracy is improved, but the amount of additional data required increases significantly

Engineering Contradiction:
Improve3D reconstruction accuracyVSAvoiddata volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

Instead of creating dense depth maps or optical flow maps that require storing data for every pixel, the patent captures multiple discrete 2D images from different viewpoints. These images serve as copies of the scene from various angles, allowing 3D reconstruction through geometric relationships between the copies rather than through dense per-pixel depth information, thereby maintaining accuracy while reducing data volume

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If computer generation of polygons or texture mapping is used to create 3D models, then 3D model quality is improved, but processing times and resource requirements increase

Engineering Contradiction:
Improve3D model qualityVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent performs view synthesis and 3D reconstruction operations in advance during the content creation phase, generating the multi-view representation before playback. This preliminary processing allows the system to store pre-computed 3D data that can be efficiently transmitted and displayed without requiring real-time polygon generation or texture mapping during playback, thereby improving processing speed while maintaining model quality

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10958891B2Visual annotation using tagging sessions
Publication Date: 2021.03.23 FUSION INC
  • US10958891B2 patent drawing
  • US10958891B2 patent drawing
  • US10958891B2 patent drawing

AI summary

Various embodiments of the present invention relate generally to systems and methods for analyzing and manipulating images and video. In particular, a multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. After the MVIDMR of the object is generated, a tag can be placed at a location on the object in the MVIDMR. The locations of the tag in the frames of the MVIDMR can vary from frame to frame as the view of the object changes. When the tag is selected, media content can be output which shows details of the object at location where the tag is placed. In one embodiment, the object can be car and tags can be used to link to media content showing details of the car at the locations where the tags are placed.