Bokeh-Based Distance Estimation Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating high-accuracy statistical models for estimating distances using monocular cameras is challenging due to the difficulty in preparing large datasets, especially in varying environments, as existing methods require extensive data collection and precise distance measurements.

Innovation Solution

A learning method that utilizes a statistical model to predict bokeh values from images captured by a monocular camera, where the model is trained using multi-view images from different viewpoints, allowing for online learning and adaptation to new environments without requiring explicit distance labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a statistical model is trained using traditional methods with explicit distance labels, then measurement precision may be improved, but data preparation complexity and time increase significantly

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates virtual distance labels by copying depth information from pre-trained depth estimation models or structure-from-motion algorithms, avoiding the need for manual measurement while maintaining training data quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces an intermediary depth prediction model that generates pseudo-labels for training the bokeh-based distance estimation model, bridging the gap between available data and required training labels

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a statistical model is trained on domain-specific data, then measurement precision improves for that domain, but adaptability to new environments deteriorates

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent trains the statistical model on multi-domain data including indoor and outdoor environments, making the model universally applicable across different settings rather than specialized for a single domain

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent incorporates domain adaptation techniques that adjust model parameters based on environmental characteristics, allowing the model to adapt to new domains while maintaining core learning

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If extensive distance measurements are collected for training, then measurement precision improves, but device complexity and measurement difficulty increase

Engineering Contradiction:
Improvedistance measurement accuracyVSAvoiddistance measurement difficulty
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent enables the system to generate its own training labels automatically through computational methods, eliminating the need for external measurement devices or manual annotation processes

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11651504B2Learning method, storage medium and image processing device
Publication Date: 2023.05.16 KK TOSHIBA
  • US11651504B2 patent drawing
  • US11651504B2 patent drawing
  • US11651504B2 patent drawing

AI summary

According to one embodiment, a learning method for causing a statistical model to learn is provided. The statistical model is generated by learning a bokeh caused in a first image captured in a first domain in accordance with a distance to a first subject included in the first image, the method includes acquiring a plurality of second images by capturing a second subject from multiple viewpoints in a second domain other than the first domain, and causing the statistical model to learn using the second images.