Multi-View Video Processing with Spatial Region Encapsulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently transmitting video resources for multi-view videos, especially when multiple view groups are involved, leading to inefficiencies in processing and user experience.

Innovation Solution

A method and apparatus for processing multi-view video data that involves acquiring and dividing the data into view groups, determining spatial region information for each group, and encapsulating this information for efficient transmission and playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If all view groups are transmitted to the user, then the user can access any view group, but the transmission efficiency decreases and unnecessary data transfer increases

Engineering Contradiction:
Improveuser access to view groupsVSAvoidtransmission efficiency
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent segments the multi-view video data into multiple view groups, each with independent spatial region information. This allows the system to transmit only the necessary view groups to the user based on their selection, rather than transmitting all view groups, thereby improving transmission efficiency while maintaining user access flexibility

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts spatial region information from each view group and transmits this metadata separately. This extraction enables the user to quickly identify and select desired view groups without receiving the complete video data, reducing unnecessary data transfer and improving transmission efficiency

Inventive Principle:
Principle #2Taking out (Extraction)

2Speed

If spatial region information is added to each view group, then user selection speed improves, but the data structure complexity increases

Engineering Contradiction:
Improveuser selection speedVSAvoiddata structure complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent performs preliminary organization of view groups with their spatial region information before transmission. By pre-structuring the data with clear spatial region metadata, the system enables fast user selection without requiring complex real-time processing, thus improving selection speed while keeping the data structure manageable

Inventive Principle:
Principle #10Preliminary action

3Productivity

If view group division is performed, then media resource transmission becomes targeted, but the processing complexity increases

Engineering Contradiction:
Improvemedia resource transmission efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the multi-view video into distinct view groups with independent spatial region information. This segmentation enables targeted media resource transmission by allowing the system to identify and transmit only the specific view groups needed, improving transmission efficiency while distributing the processing complexity across manageable units

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12323569B2Method and apparatus for processing multi-view video with region information from view group
Publication Date: 2025.06.03 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12323569B2 patent drawing
  • US12323569B2 patent drawing
  • US12323569B2 patent drawing

AI summary

A computer device acquires multi-view video data that includes video data of multiple views. The computer device performs view group division on the multi-view video data based on the multiple views to obtain at least one view group. The computer device determines first spatial region information of the at least one view group. The first spatial region information includes information of a three-dimensional spatial region where the at least one view group is located. The computer device encapsulates the multi-view video data and the first spatial region information.