3D Video Encoding with Metadata-Based Key Frame Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for encoding 3D videos struggle to efficiently reduce data amounts while maintaining image quality, particularly in setting appropriate reference frames for inter-frame prediction encoding.

Innovation Solution

An image processing apparatus and method that separately encode 3D data and texture information using inter-frame prediction, with metadata-based selection of key frames for each component to optimize encoding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If inter-frame prediction encoding is used to reduce data amount, then data compression efficiency is improved, but image quality degradation increases due to inappropriate key frame selection

Engineering Contradiction:
Improvedata amountVSAvoidimage quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments the encoding process into separate handling of 3D data and texture information, with independent key frame selection for each component based on their respective metadata characteristics. This allows optimized compression for each data type while maintaining quality where needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter selection criteria for key frames based on metadata analysis. By evaluating metadata characteristics and dynamically selecting key frames where metadata changes exceed predetermined thresholds, the system adapts the encoding parameters to actual content requirements, reducing unnecessary key frames while maintaining quality.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If key frames are frequently set to maintain image quality, then image quality is preserved, but data amount reduction efficiency decreases

Engineering Contradiction:
Improveimage qualityVSAvoiddata amount
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent dynamically changes the key frame selection parameters based on metadata analysis. By setting key frames only when metadata changes exceed predetermined thresholds, the system reduces the frequency of key frames compared to traditional methods, thereby reducing data amount while maintaining quality only when necessary.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses its own metadata to automatically determine key frame placement without external intervention or overly frequent reference frames. The metadata-driven approach allows the encoding process to self-regulate key frame frequency based on actual content requirements.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If separate encoding of 3D data and texture information is implemented, then encoding precision is improved, but device complexity increases

Engineering Contradiction:
Improveencoding precisionVSAvoidencoding complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the encoding process into separate streams for 3D data and texture information, each with independent key frame selection based on their respective metadata. This segmentation improves encoding precision by allowing component-specific optimization while managing complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250193449A1Image processing apparatus and image processing method
Publication Date: 2025.06.12 CANON KK
  • US20250193449A1 patent drawing
  • US20250193449A1 patent drawing
  • US20250193449A1 patent drawing

AI summary

Disclosed is an image processing apparatus that reduces the data amount of a 3D video by use of the corrections between frames. The image processing apparatus obtains 3D video data each frame of which includes 3D data and texture information. The apparatus performs encodes the 3D video data by separately performing inter-frame prediction encoding of the 3D data and inter-frame prediction encoding of the texture information. The apparatus, based on metadata of each frame, separately selects a key frame for performing the inter-frame prediction encoding of the 3D data and a key frame for performing the inter-frame prediction encoding of the texture information.