Image Processing Apparatus Selecting Cameras for Virtual Viewpoint Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing techniques require significant time for machine learning due to the need to associate captured images from multiple cameras with virtual viewpoint images, leading to inefficiencies in the learning process.

Innovation Solution

An image processing apparatus that selects a real camera based on virtual viewpoint information to generate a learning dataset by associating a captured image with a learning virtual viewpoint image, reducing the number of cameras involved in machine learning and optimizing the learning process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning is performed by associating captured images from all plurality of cameras with corresponding virtual viewpoint images, then image quality improvement is achieved, but learning time becomes excessively long

Engineering Contradiction:
Improveimage qualityVSAvoidlearning time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the necessary subset of camera images for machine learning by selecting cameras whose optical axes are close to the virtual camera's optical axis at each viewpoint. This extraction principle reduces the learning dataset from all camera images to only the essential ones, significantly decreasing learning time while maintaining image quality improvement effectiveness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by treating different cameras differently based on their spatial relationship with the virtual camera. Instead of uniform processing of all cameras, the system selectively includes only those cameras with optical axes close to the virtual camera's optical axis, creating a localized learning approach that focuses on relevant data only.

Inventive Principle:
Principle #3Local quality

2Reliability

If all plurality of cameras are used for machine learning, then comprehensive training data is obtained, but device complexity and processing load increase

Engineering Contradiction:
Improvetraining data comprehensivenessVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary subset of camera images for machine learning by selecting cameras whose optical axes are close to the virtual camera's optical axis at each viewpoint. This extraction principle reduces the learning dataset from all camera images to only the essential ones, significantly decreasing learning time while maintaining image quality improvement effectiveness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by using only a subset of cameras for learning rather than all cameras. The system determines, for each viewpoint, which cameras have optical axes close to the virtual camera's optical axis and uses only those cameras' images for learning, avoiding the excessive processing load of handling all camera data while still obtaining sufficient training data.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240323331A1Image processing apparatus, image processing method, and storage medium
Publication Date: 2024.09.26 CANON KK
  • US20240323331A1 patent drawing
  • US20240323331A1 patent drawing
  • US20240323331A1 patent drawing

AI summary

To make it possible to obtain a highly accurate trained model with a small amount of learning time. A plurality of captured images corresponding to each of a plurality of imaging devices and virtual viewpoint information predefining a virtual viewpoint for generating a virtual viewpoint image are obtained and based on the virtual viewpoint information, an imaging device that is referred to in machine learning is selected from among the plurality of imaging devices. Then, the machine learning is performed by using learning dataset associating a captured image of a selected imaging device as reference data with a learning virtual viewpoint image generated by taking a viewpoint of the imaging device as a virtual viewpoint.