Encoding Free Viewpoint Data in Movie Containers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating real-time, immersive augmented reality experiences with virtual objects in real-world scenes are hindered by the difficulty in preprocessing and rendering 3D items within real-time video streams, particularly due to limitations in available computing resources and the need for precise coordinate translation between real-world and virtual scenes.

Innovation Solution

A method of encoding arbitrary data, such as free-view data defining 3D models, into a standard movie data container like an MPEG container, allowing for the insertion of non-video data into video streams with predefined labels, enabling efficient processing and rendering of 3D models with varying frame rates and resolutions within the same video stream or separate streams.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If non-video data (free-view data) is inserted into video streams for efficient processing, then the portability and wearability of AR systems is improved, but the complexity of data container structure increases

Engineering Contradiction:
Improveportability and wearabilityVSAvoiddata container structure
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent combines non-video free-view data with video streams by inserting them into the same movie data container (MP4 format). This merging approach allows both video and 3D model data to be processed together in the video stream, enabling efficient real-time rendering on portable devices with limited resources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses a universal video stream format (MP4 container) to carry multiple types of data - both traditional video frames and non-video free-view 3D model data. This multi-functional approach allows a single data container to serve multiple purposes, reducing the need for separate processing pipelines and improving system portability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If free-view data is encoded into standard movie containers for efficient processing, then real-time rendering performance is improved, but the precision of coordinate translation between real-world and virtual scenes may be compromised

Engineering Contradiction:
Improvereal-time rendering performanceVSAvoidcoordinate translation precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces a coordinate translation mechanism that acts as an intermediary between the real-world camera coordinates and the virtual 3D model coordinates. This mediator ensures accurate spatial alignment while allowing the data to be processed through the efficient video rendering pipeline, maintaining both precision and performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms free-view 3D model data into a format compatible with video stream encoding by changing its parameter representation. The 3D models are encoded with metadata including position, orientation, and scale parameters that enable accurate coordinate translation while being processed through standard video decoding pipelines for real-time performance.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If multiple video streams with different frame rates and resolutions are used for 3D models, then the quality of virtual object rendering is improved, but the device complexity increases

Engineering Contradiction:
Improvevirtual object rendering qualityVSAvoidstream management complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments different 3D model data into separate video streams with optimized frame rates and resolutions based on their specific requirements. This segmentation allows each virtual object to be rendered at appropriate quality levels while maintaining manageable stream complexity through standardized container formatting.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10306254B2Encoding free view point data in movie data container
Publication Date: 2019.05.28 SEIKO EPSON CORP
  • US10306254B2 patent drawing
  • US10306254B2 patent drawing
  • US10306254B2 patent drawing

AI summary

Multiple Holocam Orbs observe a real-life environment and generate an artificial reality representation of the real-life environment. Depth image data is cleansed of error due to LED shadow by identifying the edge of a foreground object in an (near infrared light) intensity image, identifying an edge in a depth image, and taking the difference between the start of both edges. Depth data error due to parallax is identified noting when associated text data in a given pixel row that is progressing in a given row direction (left-to-right or right-to-left) reverses order. Sound sources are identified by comparing results of a blind audio source localization algorithm, with the spatial 3D model provided by the Holocam Orb. Sound sources that corresponding to identifying 3D objects are associated together. Additionally, types of data supported by a standard movie data container, such as an MPEG container, is expanding to incorporate free viewpoint data (FVD) model data. This is done by inserting FVD data of different individual 3D objects at different sample rates into a single video stream. Each 3D object is separately identified by a separately assigned ID.