Texture-Based Immersive Video Coding Without Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current immersive video coding standards, such as MPEG Immersive Video (MIV), face inefficiencies in bandwidth utilization and fail to fully exploit angular redundancy across views, leading to incomplete immersive cues due to the reliance on depth information and lack of consideration for textural information in encoding and decoding processes.

Innovation Solution

The proposed solution involves a texture-based immersive video coding approach that excludes depth information from the bitstream, utilizing only texture data per input view and inferring depth from coded texture videos, allowing for more efficient bandwidth usage and improved delivery of immersive cues by identifying correspondence patches across input views and transmitting only texture content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth information is included in the bitstream for immersive video coding, then reconstruction quality is improved, but bandwidth utilization deteriorates due to redundant data transmission

Engineering Contradiction:
Improvereconstruction qualityVSAvoidbandwidth utilization
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent extracts and transmits only texture information from the multi-view video data, separating it from depth information. By taking out only the essential texture components and excluding redundant depth data, the system achieves efficient bandwidth utilization while maintaining reconstruction quality through texture-based depth inference at the decoder side.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of the conventional approach of encoding and transmitting depth information to achieve 3D reconstruction, the patent inverts the approach by transmitting only texture information and inferring depth at the decoder. This inversion eliminates the need for depth encoding while preserving reconstruction quality through texture-based inference.

Inventive Principle:
Principle #13The other way round (Inversion)

2Loss of information

If all views are encoded separately in current immersive video standards, then complete immersive cues are provided, but bandwidth consumption increases due to failure to exploit angular redundancy

Engineering Contradiction:
Improveimmersive cuesVSAvoidbandwidth consumption
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The patent merges multiple view textures into a single unified texture representation that captures all essential immersive cues. By combining information from multiple views into one integrated texture stream, the system eliminates redundant data transmission while preserving complete immersive information for 3D reconstruction.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The transmitted texture information serves multiple functions simultaneously: it provides visual content for display, enables depth inference for 3D reconstruction, and contains angular redundancy exploitation for efficient compression. This multi-functionality allows a single texture stream to replace multiple separate view encodings.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12069302B2Texture based immersive video coding
Publication Date: 2024.08.20 INTEL CORP
  • US12069302B2 patent drawing
  • US12069302B2 patent drawing
  • US12069302B2 patent drawing

AI summary

Methods, apparatus, systems and articles of manufacture for texture based immersive video coding are disclosed. An example apparatus includes a correspondence labeler to (i) identify first unique pixels and first corresponding pixels included in a plurality of pixels of a first view and (ii) identify second unique pixels and second corresponding pixels included in a plurality of pixels of a second view; a correspondence patch packer to (i) compare adjacent pixels in the first view and (ii) identify a first patch of unique pixels and a second patch of corresponding pixels based on the comparison of the adjacent pixels and the correspondence relationships, the second patch of corresponding pixels tagged with a correspondence list identifying corresponding patches in the second view; and an atlas generator to generate at least one atlas to include in encoded video data, the encoded video data not including depth maps.