Multi-View Video Synthesis Depth Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current immersive video processing schemes face challenges in efficiently transmitting and synthesizing intermediate views, particularly due to high bit rates and poor quality depth information, which can lead to motion sickness and increased computational load on terminals like smartphones.

Innovation Solution

A method for processing multi-view video data that allows for flexible selection of modes for obtaining synthesis data, either by decoding it from the encoded data stream or estimating it on the client side, optimizing encoding cost and synthesis quality based on available tools and content characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If depth maps are calculated and transmitted for all views, then the quality of intermediate view synthesis is improved, but the bit rate and transmission cost increase significantly

Engineering Contradiction:
Improvesynthesis qualityVSAvoidbit rate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts and transmits only the necessary depth information for intermediate view synthesis rather than complete depth maps for all views. The depth estimator identifies and transmits only the depth data that will be useful for synthesizing intermediate views, eliminating redundant depth information transmission.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The depth estimator on the server side performs preliminary analysis to determine which depth information will be needed for intermediate view synthesis before encoding. This allows the system to prepare and transmit only the relevant depth data in advance, avoiding the need to transmit all depth maps.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If complete depth maps are generated and transmitted, then synthesis quality is improved, but redundant information increases transmission cost

Engineering Contradiction:
Improvedepth map qualityVSAvoidtransmission cost
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The system extracts only the essential depth information needed for intermediate view synthesis from the complete depth maps. The depth estimator identifies regions and depth data that are actually required for synthesis, separating useful information from redundant data before transmission.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different quality levels to different regions of depth maps based on their importance for intermediate view synthesis. Critical regions are transmitted with high quality, while less important regions are either transmitted with lower quality or omitted entirely, optimizing the balance between synthesis quality and transmission cost.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If depth information is captured with dedicated sensors, then depth quality is improved, but device complexity and cost increase

Engineering Contradiction:
Improvedepth information qualityVSAvoidacquisition system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces dedicated depth sensing hardware with computational depth estimation methods. Instead of using complex sensor systems to capture depth information, the system uses software-based depth estimation algorithms that analyze texture images to generate depth maps, significantly reducing device complexity while maintaining acceptable depth quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The depth estimator acts as an intermediary that converts readily available texture image data into depth information. This intermediary process allows the system to obtain depth data without direct sensor measurement, using computational methods to bridge the gap between 2D images and 3D depth information.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If multiple captured views are transmitted, then intermediate view synthesis capability is improved, but the number of views and data amount increase

Engineering Contradiction:
Improveview synthesis capabilityVSAvoiddata amount
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system extracts and transmits only the essential view data and associated depth information needed for intermediate view synthesis. Rather than transmitting all captured views, the depth estimator identifies the minimum set of view data required to reconstruct intermediate views, reducing the overall data amount while maintaining synthesis capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The depth estimator performs preliminary analysis to determine which captured views contain the most valuable information for intermediate view synthesis. This allows the system to select and transmit only the most useful views in advance, reducing data transmission requirements while preserving synthesis flexibility.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12284385B2Method and device for processing multi-view video data with synthesis data
Publication Date: 2025.04.22 ORANGE SA
  • US12284385B2 patent drawing
  • US12284385B2 patent drawing
  • US12284385B2 patent drawing

AI summary

A method for processing multi-view video data including, for at least one block of an image of a view encoded in an encoded data stream representing the multi-view video, obtaining at least one information item, which specifies a mode for obtaining a synthesis data item, from among first and second obtaining modes. The synthesis data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view not being encoded in the encoded data stream. The first obtaining mode involves decoding an information item representing the synthesis data item from the encoded data stream, the second obtaining mode involves obtaining the synthesis data item from at least the reconstructed encoded image. At least one part of an image of the intermediate view is synthesised from at least the reconstructed encoded image and the synthesis data item obtained by the specified mode.