Multi-view Video Synthesis via Extracted Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current immersive video processing systems face challenges in efficiently transmitting and synthesizing intermediate views, particularly due to high bit rates and the need for accurate depth maps, which are often of poor quality and redundant.

Innovation Solution

The method involves synthesizing intermediate views on the client side using data obtained from decoded and reconstructed views, eliminating the need for transmitting depth maps and reducing the encoding rate. This is achieved by extracting synthesis data, such as depth maps, on the decoder side from decoded textures, and optionally refining this data using neural networks or refinement data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If depth maps are calculated and transmitted prior to encoding, then view synthesis quality is improved, but transmission bit rate increases significantly

Engineering Contradiction:
Improveview synthesis qualityVSAvoidtransmission bit rate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by calculating depth maps at the encoder side before encoding and transmitting them to the decoder. This allows the depth information to be prepared in advance with high quality, enabling accurate view synthesis at the decoder without requiring high bit rates for depth transmission, since the depth maps are already computed and optimized before transmission.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts depth information from the multi-view video data at the encoder side using depth estimation algorithms. By separating the depth map extraction and transmission from the texture transmission, the system can optimize depth quality independently while controlling the overall bit rate, as depth maps constitute a significant but manageable portion of the total data.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If complete depth maps are generated and transmitted, then view synthesis coverage is improved, but redundant information increases transmission cost

Engineering Contradiction:
Improveview synthesis coverageVSAvoidtransmission data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by transmitting complete depth maps only when necessary, and using selective depth transmission based on the specific view synthesis requirements. The encoder can determine which depth regions are actually needed for the requested intermediate views and transmit only those relevant portions, reducing redundant data while maintaining synthesis coverage where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses partial action by transmitting depth maps at a resolution and detail level that is sufficient for the specific synthesis task, rather than always transmitting complete high-resolution depth maps. The encoder can adjust the depth map transmission quality and coverage based on the required synthesis accuracy and available bit rate, avoiding excessive data transmission.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If depth estimation is performed on server side, then synthesis data availability is improved, but encoding complexity increases

Engineering Contradiction:
Improvesynthesis data availabilityVSAvoidencoding complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies self-service by having the encoder perform depth estimation automatically as part of the encoding process. The depth maps are generated by the encoding system itself using integrated depth estimation algorithms, eliminating the need for separate external depth calculation steps and making the synthesis data self-sufficient within the encoded bitstream.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent merges the depth estimation function with the video encoding process. The depth map generation, encoding, and transmission are combined into a unified encoding pipeline, reducing overall system complexity by eliminating separate processing stages and integrating depth synthesis capabilities directly into the encoder.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If multiple captured views are transmitted, then view synthesis accuracy is improved, but transmission rate increases

Engineering Contradiction:
Improveview synthesis accuracyVSAvoidtransmission efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent uses depth maps as an intermediary element that enables accurate view synthesis from fewer captured views. Instead of transmitting multiple high-resolution captured views, the system transmits captured views along with computed depth maps, which act as intermediary data that allows the decoder to synthesize additional views with accuracy comparable to having multiple captured views, but with lower transmission requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12278937B2Method and device for processing multi-view video data
Publication Date: 2025.04.15 ORANGE SA
  • US12278937B2 patent drawing
  • US12278937B2 patent drawing
  • US12278937B2 patent drawing

AI summary

A method and a device for processing multi-view video data. The multi-view video data includes at least one part of a decoded image of at least one view of the multi-view video, from an encoded data stream representative of the multi-view video. At least one item of data, referred to as synthesis data, is obtained from at least the one part of the decoded image, and at least one image of an intermediate view of the multi-view video not encoded in the encoded data stream is synthesized from at least the one part of the decoded image and from the synthesis data obtained.