Multiview Video Random Access via V-Frame Insertion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multiview video encoding and decoding methods fail to efficiently utilize temporal and spatial correlations between views, leading to increased network bandwidth and storage requirements due to independent encoding of each camera view, and lack a method to account for correlations in time.

Innovation Solution

A method that decomposes a multiview bit stream into low and high band frames using motion-compensated temporal filtering (MCTF) and disparity-compensated inter-view filtering (DCVF), with a reference picture list managing temporal, spatial, and synthesized reference pictures for prediction, and inserting V-frames for random temporal access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If independent encoding is used for each camera view, then encoding simplicity is maintained, but compression efficiency decreases and network bandwidth and storage increase

Engineering Contradiction:
Improveencoding simplicityVSAvoidnetwork bandwidth
Core Design Contradiction:
Ease of manufactureVSLoss of energy

Solution Approach 1:

The patent merges the encoding of multiple camera views into a single integrated encoding process that exploits temporal and spatial correlations between views. Instead of independently encoding each view, the system uses prediction models that reference other views and time points, thereby reducing the total data量 required for transmission and storage while maintaining encoding feasibility through standardized prediction techniques.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of manufacture

If independent encoding is used for each camera view, then encoding simplicity is maintained, but compression efficiency decreases and storage requirements increase

Engineering Contradiction:
Improveencoding simplicityVSAvoidstorage requirements
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent merges the encoding of multiple camera views into a single integrated encoding process that exploits temporal and spatial correlations between views. Instead of independently encoding each view, the system uses prediction models that reference other views and time points, thereby reducing the total data量 required for transmission and storage while maintaining encoding feasibility through standardized prediction techniques.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If conventional prediction methods are used, then encoding compatibility is maintained, but temporal and spatial correlations between views are not utilized

Engineering Contradiction:
Improveencoding compatibilityVSAvoidcompression efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements a universal prediction framework that can operate within existing video coding standards while simultaneously exploiting temporal and spatial correlations. The prediction model is designed to be compatible with conventional encoders but enhances compression efficiency by utilizing multiple reference sources (temporal references from previous frames and spatial references from other camera views) in a multi-functional prediction process.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of operation

If no prediction dependency information is provided, then decoding simplicity is maintained, but random access efficiency decreases

Engineering Contradiction:
Improvedecoding simplicityVSAvoidrandom access time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-calculating and storing prediction dependency information during the encoding process. This metadata about which frames and views are needed for prediction is prepared in advance and transmitted with the encoded data, enabling the decoder to efficiently determine the exact sequence of decoding operations required for random access without needing to analyze the entire bitstream in real-time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7903737B2Method and system for randomly accessing multiview videos with known prediction dependency
Publication Date: 2011.03.08 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US7903737B2 patent drawing
  • US7903737B2 patent drawing
  • US7903737B2 patent drawing

AI summary

A method randomly accesses multiview videos. Multiview videos are acquired of a scene with corresponding cameras arranged at poses, such that there is view overlap between any pair of cameras. V-frames are generated from the multiview videos. The V-frames are encoded using only spatial prediction. Then, the V-frames are inserted periodically in an encoded bit stream to provide random temporal access to the multiview videos. Additional view dependency information enables the decoding of a reduced number of frames prior to accessing randomly a target frame for a specified view and time, and decoding the target frame.