Multi-view Video Processing Atlas Patch Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current immersive media technologies supporting 3DOF+ experiences face inefficiencies in media content transmission and rendering due to redundant data from multi-camera setups, lacking effective methods for representing and rendering media content that accounts for user movement within a limited range.

Innovation Solution

The method involves constructing media content by extracting and synthesizing patches from atlases based on user viewing position, direction, and viewport, using a hierarchical structure that combines texture and depth components from multiple views, allowing for efficient rendering of three-dimensional stereoscopic video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multi-camera setups are used to capture immersive media content, then the coverage and viewing experience are improved, but redundant data increases transmission and storage requirements

Engineering Contradiction:
Improveviewing experienceVSAvoiddata volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent segments the immersive media content into multiple atlases, where each atlas corresponds to a specific view captured by a camera. This segmentation allows the system to organize and manage multi-camera data efficiently, transmitting only the necessary atlas data for the user's current viewpoint rather than all camera data, thereby reducing redundant data transmission while maintaining comprehensive viewing coverage.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If all views from multiple cameras are transmitted, then complete visual information is provided, but transmission efficiency decreases

Engineering Contradiction:
Improvevisual information completenessVSAvoidtransmission efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts and transmits only the essential view information corresponding to the user's current viewpoint and viewport. By identifying which atlas data is necessary for the user's present viewing state and transmitting only that subset, the system maintains complete visual information for the active viewport while significantly improving transmission efficiency by eliminating redundant data from other viewpoints.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If traditional rendering methods are used for multi-view content, then all views can be displayed, but rendering complexity and computational load increase

Engineering Contradiction:
Improveview rendering capabilityVSAvoidrendering complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary organization of multi-camera data into structured atlases during content creation or preprocessing. Each atlas is pre-configured with view-specific information including camera parameters, projection data, and spatial relationships. This preliminary structuring enables the rendering system to efficiently retrieve and display only the necessary atlas data for the user's current viewpoint, reducing real-time rendering complexity while maintaining full multi-view capability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12015756B2Multi-view video processing method and apparatus
Publication Date: 2024.06.18 ZTE CORP
  • US12015756B2 patent drawing
  • US12015756B2 patent drawing
  • US12015756B2 patent drawing

AI summary

Methods, apparatus, and systems for effectively reducing media content transmissions and efficiently rendering immersive media contents are disclosed. In one example aspect, a method includes requesting, by a user, media files from a server according to the current viewing position and viewing direction of the user, receiving, by the user, the media files from the server according to the current viewing position and the viewing direction of the user, extracting a patch of an atlas, and synthesizing the visual content in the current window area of the user, and obtaining, by the user, three-dimensional stereoscopic video content according to the current viewing position and viewing direction of the user.