Depth Segmentation Patches for Real-Time Multi-View Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating multi-view videos require significant processing power due to complex depth estimation, especially in real-time applications like live sporting events, leading to inefficiencies in creating depth and texture atlas data.

Innovation Solution

Segment foreground objects from source view images and depth maps into patches, generating patch texture, depth, and transparency maps, using background depth maps for robust segmentation, and constructing an atlas with these patches to reduce processing requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If complex depth estimation algorithms are used to generate multi-view videos, then rendering quality is improved, but processing power requirements increase significantly

Engineering Contradiction:
Improverendering qualityVSAvoidprocessing power
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The patent segments the image into multiple depth layers, where each layer contains pixels with similar depth values. This segmentation allows the system to process and render only the necessary layers for a given viewpoint, reducing the overall processing power required while maintaining rendering quality through selective processing of depth information.

Inventive Principle:
Principle #1Segmentation

2Productivity

If depth and texture atlas data are created in real-time for broadcast applications, then productivity is improved, but processing power requirements increase

Engineering Contradiction:
Improvereal-time generation capabilityVSAvoidprocessing power
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

By segmenting the scene into depth layers and processing only the necessary layers for real-time broadcast, the system achieves real-time productivity without requiring full processing power for the entire scene. This selective processing enables real-time atlas data creation while managing processing resources efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by processing only the necessary depth layers and image regions required for the current viewpoint and broadcast requirements, rather than processing the entire scene. This partial processing approach enables real-time productivity by avoiding unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If layered formats like LDI, MPI, or MSI are used to handle occlusion, then viewing zone is improved, but ease of manufacture decreases

Engineering Contradiction:
Improveviewing zoneVSAvoiddifficulty to produce
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent uses segmentation to divide the image into depth layers, which simplifies the handling of occlusion compared to traditional layered formats. Each depth layer can be independently processed and blended, making the system easier to manufacture and implement while still providing comprehensive viewing coverage.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12430836B2Depth segmentation in multi-view videos
Publication Date: 2025.09.30 KONINKLIJKE PHILIPS NV
  • US12430836B2 patent drawing
  • US12430836B2 patent drawing
  • US12430836B2 patent drawing

AI summary

A method of depth segmentation for the generation of a multi-view video data. The method comprises obtaining a plurality of source view images and source view depth maps representative of a 3D scene from a plurality of sensors. Foreground objects in the 3D scene are segmented from the source view images and/or the source view depth maps. One or more patches are then generated for each source view image and source view depth map containing at least one foreground object, wherein each patch corresponds to a foreground object and wherein generating a patch comprises generating a patch texture image, a patch depth map and a patch transparency map based on the source view images and the source view depth maps.