Multi-viewpoint Image Encoding Reference Viewpoint Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image encoding techniques for multiple viewpoints struggle with high data volume, especially in moving images, due to inefficient inter-viewpoint prediction methods, leading to increased code amounts and difficulties in storage and transmission.

Innovation Solution

An image processing apparatus and method that selects a reference viewpoint based on capturing conditions and performs predicting coding by referencing only the frame image from the reference viewpoint, without considering other viewpoints, to optimize encoding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If inter-viewpoint prediction is performed using the central viewpoint as reference, then encoding between viewpoints is enabled, but the prediction residual becomes large when exposure times differ greatly, increasing the code amount

Engineering Contradiction:
Improveencoding efficiencyVSAvoidcode amount
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent dynamically changes the reference viewpoint selection based on capturing conditions (exposure time, aperture, focal length). Instead of always using the central viewpoint, the system selects the viewpoint with capturing conditions closest to the target viewpoint as the reference, thereby minimizing prediction residuals and code amount while maintaining encoding efficiency.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If encoding is performed for each viewpoint independently, then encoding simplicity is maintained, but the code amount increases due to lack of inter-viewpoint prediction

Engineering Contradiction:
Improveencoding complexityVSAvoidcode amount
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by performing inter-viewpoint prediction selectively rather than universally. It identifies viewpoints with similar capturing conditions and performs prediction only for those, while encoding other viewpoints independently. This localized approach reduces code amount without significantly increasing encoding complexity.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If a plurality of captured images are composited into HDR image after capturing, then user flexibility is improved, but the data amount becomes large and difficult to store or transmit

Engineering Contradiction:
Improveuser flexibilityVSAvoiddata amount
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent performs encoding and compression of captured images from multiple viewpoints before storage or transmission. By applying inter-viewpoint prediction and compressing the data in advance, it reduces the data amount that needs to be stored or transmitted, while still preserving the capability to composite HDR images and meet user requests later.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9363432B2Image processing apparatus and image processing method
Publication Date: 2016.06.07 CANON KK
  • US9363432B2 patent drawing
  • US9363432B2 patent drawing
  • US9363432B2 patent drawing

AI summary

For each of respective viewpoints, the capturing condition at the viewpoint, and each frame image captured from the viewpoint in accordance with the capturing condition are acquired. One of the viewpoints is selected as a reference viewpoint by using the acquired capturing conditions. Predicting coding and intra-coding are performed for the acquired images. For each of the respective viewpoints for each frame, the coding result of the frame image captured from the viewpoint by predicting coding or intra-coding is output. When performing predicting coding for an image captured from the reference viewpoint, predicting coding is performed by referring to the image captured from the reference viewpoint without referring to images captured from the viewpoints other than the reference viewpoint.