Facet-Based Video Encoding for Resource-Limited VR Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for stitching together wide field of view content from multiple images require significant computing resources, making it difficult for users to capture and share panoramic and virtual reality content on resource-limited devices such as smartphones or head-mounted displays.
Innovation Solution
A system that encodes panoramic image content by partitioning images into facets, transforming, and encoding them using operations like rotation, flipping, and scaling, allowing for efficient encoding and transmission on resource-limited devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of moving object
If multiple cameras are used to capture wide field of view content, then the field of view coverage is improved, but the computing resources required for stitching increase significantly
Solution Approach 1:
The patent divides the panoramic image into multiple cubic facets (six faces of a cube), allowing each facet to be processed and encoded independently. This segmentation reduces the computational complexity of stitching by treating each facet as a separate unit that can be encoded using standard video coding techniques, rather than processing the entire high-resolution panoramic image as a single unit.
2Manufacturing precision
If high resolution panoramic images are processed, then the image quality is improved, but the memory resources consumed increase significantly
Solution Approach 1:
By segmenting the high-resolution panoramic image into six smaller cubic facets, the patent reduces the memory footprint required for processing. Each facet can be loaded, processed, and discarded independently, avoiding the need to hold the entire high-resolution image in memory simultaneously. This enables processing of high-quality images on devices with limited memory resources.
3Ease of operation
If image stitching is performed on portable devices, then the convenience of content creation is improved, but the processing power required exceeds available resources
Solution Approach 1:
The patent enables portable devices to perform stitching by dividing the complex stitching operation into simpler facet-based encoding operations. Each facet can be processed using efficient video coding algorithms that are less computationally intensive than traditional panoramic stitching methods, making the operation feasible on mobile devices with limited processing power.
Solution Approach 2:
The patent uses reference facets to predict and encode target facets, reducing the computational burden. By copying and transforming reference facet data (through rotation, flipping, and scaling operations) to generate predictions for adjacent facets, the system avoids performing full stitching calculations for every facet, significantly reducing processing requirements.
4Measurement precision
If traditional stitching algorithms are used, then the stitching accuracy is improved, but the computational expense increases
Solution Approach 1:
The patent achieves efficient stitching by copying reference facet information and applying geometric transformations (rotation, flipping, scaling) to generate predictions for target facets. This copying approach maintains stitching accuracy by preserving the original image data while reducing computational expense through efficient transformation operations rather than complex alignment algorithms.
Solution Approach 2:
The patent replaces traditional mechanical stitching algorithms with a facet-based encoding approach that uses standard video coding techniques. Instead of performing complex image registration and blending operations, the system encodes each facet independently using transform coding and prediction methods, achieving comparable stitching results with significantly reduced computational complexity.
Data Source
AI summary
A device includes a processor that is configured to obtain first facets of a first wide field-of-view image. An object is identified in a facet of the first facets. A second wide field-of-view image is obtained. A location of the object is identified in the second wide field-of-view image. Using the location of the object, the second wide field-of-view image is partitioned into second facets such that no boundary of any of the second facets overlaps the object. The second facets are then encoded in a compressed bitstream.


