Additional View Block Packing for 3DoF+ Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing coding schemes for immersive media, particularly 3DoF+ video, face challenges in efficiently reducing redundancy and computational complexity for multi-view image or video data, especially in virtual reality applications, where high redundancy exists among views captured from slightly different positions, leading to increased bandwidth and computational demands.

Innovation Solution

A method involving block rearrangement and transformation of additional views to create a packed additional view, which is then split into parts and transformed to reduce size, followed by encoding with metadata describing the process, using lossless compression for metadata and lossy compression for video data, to enhance coding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple views are captured from slightly different positions to enable 3DoF+ video, then immersive experience is improved, but redundancy among views increases

Engineering Contradiction:
Improveimmersive experienceVSAvoidredundancy
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent combines multiple additional views into a single packed additional view by identifying and merging common image data regions across different views. This merging process eliminates redundancy while preserving the immersive experience by maintaining all unique visual information from multiple camera positions.

Inventive Principle:
Principle #5Merging (Combining)

2Manufacturing precision

If all views are encoded separately to maintain quality, then image quality is preserved, but bitrate increases

Engineering Contradiction:
Improveimage qualityVSAvoidbitrate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent merges multiple additional views into a single packed additional view that contains all unique visual information from the original views. This consolidated representation significantly reduces the bitrate required to encode the same visual content while maintaining image quality, as common regions are encoded only once rather than separately in each view.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of manufacture

If conventional video codecs are used directly on multi-view data, then encoding simplicity is maintained, but coding efficiency is insufficient

Engineering Contradiction:
Improveencoding simplicityVSAvoidcoding efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent performs preliminary packing of multiple additional views into a single consolidated view before encoding. This preliminary action reorganizes the multi-view data into a format that conventional video codecs can process efficiently, eliminating the need for complex multi-view encoding algorithms while significantly improving coding efficiency through reduced redundancy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4189958B1Packing of views for image or video coding
Publication Date: 2025.09.10 KONINKLIJKE PHILIPS NV
  • EP4189958B1 patent drawingFigure 1
  • EP4189958B1 patent drawingFigure 2
  • EP4189958B1 patent drawingFigure 3

AI summary

An encoder, decoder, encoding method and decoding method for 3DoF+ video are disclosed. The encoding method comprises receiving (110) multi-view image or video data comprising a basic view and at least a first additional view of a scene. The method proceeds by identifying (220) pixels in the first additional view that need to be encoded because they contain scene-content that is not visible in the basic view. The first additional view is divided (230) into a plurality of first blocks of pixels. First blocks containing at least one of the identified pixels are retained (240); and first blocks that contain none of the identified pixels are discarded. The retained blocks are rearranged (250) so that they are contiguous in at least one dimension. A packed additional view is generated (260) from the rearranged first retained blocks and encoded (264).