Statistical Model Distance Estimation Using Multi-View Bokeh Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating a high-accuracy statistical model for estimating distances to subjects using images captured by a single camera is challenging due to the need for a large dataset and the complexity of preparing such data.

Innovation Solution

A learning method that uses multi-view images captured from multiple viewpoints to cause a statistical model to learn, where the model predicts bokeh values based on the relationship between distances and bokeh occurrences, allowing for improved distance estimation without requiring extensive labeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a statistical model is trained using a large dataset with correct answer labels to improve distance estimation accuracy, then the measurement precision improves, but the ease of manufacture deteriorates due to the difficulty of preparing large labeled datasets

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoiddataset preparation difficulty
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses rendered images as synthetic copies of real-world scenes to create training data. These rendered images include ground truth depth information and simulate various camera conditions, providing a scalable source of training data without requiring manual annotation of real images.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system generates its own training data through rendering engines that automatically create image-depth pairs with known correct answers. This self-generated data eliminates the need for external data collection and manual labeling, allowing the model to train on unlimited synthetic examples.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If multi-view images are used to train the statistical model, then the adaptability improves by enabling learning without correct answer labels, but the device complexity increases due to the need for multiple capture devices or viewpoints

Engineering Contradiction:
Improvelearning capability without labelsVSAvoidmulti-view capture system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a rendered image as an intermediary between the multi-view images and the statistical model. The rendered images serve as a bridge that provides ground truth depth information, allowing the model to learn from unlabeled multi-view images through the mediation of synthetic training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The training process is segmented into separate stages: first training on rendered images with ground truth, then applying the pre-trained model to multi-view images. This segmentation allows the complex task of unsupervised learning to be broken down into manageable steps.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12243252B2Learning method, storage medium, and image processing device
Publication Date: 2025.03.04 KK TOSHIBA
  • US12243252B2 patent drawing
  • US12243252B2 patent drawing
  • US12243252B2 patent drawing

AI summary

According to one embodiment, a learning method includes acquiring first multi-view images obtained by capturing a first subject and causing a statistical model to learn, based on first and second bokeh values output from the statistical model by inputting first and second images of the first multi-view images. The causing includes acquiring a first distance from the capture device to a first subject in the first image and a second distance from the capture device to a first subject in the second image, discriminating a relationship in length between the first and second distances, and causing the statistical model to learn such that a relationship in magnitude between the first and second bokeh values is equal to the discriminated relationship.