Autofocus Eye Selection via Depth Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current autofocus systems in cameras, particularly when focusing on humans or animals, face challenges in accurately determining the eye or pupil for sharp focus, especially when the subject is not directly facing the camera, leading to inconsistent focus results.

Innovation Solution

An image processing apparatus and method that utilizes a trained model to estimate the distance between eye and nose features in an image, selecting the eye closer to the camera for focus adjustment, thereby improving autofocus accuracy by leveraging deep neural networks for feature detection and distance estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional autofocus detection methods are used, then the system is simple to implement, but the accuracy of detecting the correct eye for focus is poor when the subject is not directly facing the camera

Engineering Contradiction:
Improveeye detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from 2D image plane detection to 3D spatial reasoning by estimating depth distances from the camera to each eye. The neural network model predicts distance values for multiple eye candidates, enabling the system to select the closest eye based on inferred depth information rather than relying solely on 2D image coordinates and facial orientation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent replaces traditional mechanical/optical autofocus detection mechanisms with an AI-based neural network approach. Instead of using phase detection sensors or contrast-based methods that struggle with angled subjects, the system uses deep learning models trained on labeled data to directly predict eye positions and distances, substituting physical detection methods with computational intelligence.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If deep neural networks are used for feature detection, then the measurement precision of eye position and distance is improved, but the computational complexity and processing time increase

Engineering Contradiction:
Improveeye position and distance estimation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training neural network models on large datasets of labeled images before deployment. The models are trained offline to learn the complex mappings between image features and eye position/distance relationships, so that during actual autofocus operation, the pre-trained networks can quickly infer distances without requiring complex real-time computations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating synthetic training data through image processing and labeling. Ground truth distance maps are generated by processing labeled images through geometric transformations, creating copies of real scenarios with known answers. This allows the neural network to be trained on extensive synthetic data that mirrors real-world conditions without requiring equally extensive real-world annotated datasets.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240212193A1Image processing apparatus, method of generating trained model, image processing method, and medium
Publication Date: 2024.06.27 CANON KK
  • US20240212193A1 patent drawing
  • US20240212193A1 patent drawing
  • US20240212193A1 patent drawing

AI summary

An image processing apparatus is provided. First position estimation is performed to detect one or more parts of a first type in an image. A distance between each of the one or more parts of the first type and a part of a second type in the image is estimated. From among the one or more parts of the first type detected, one part of the first type is selected based on the distance estimated for each of the parts of the first type. Information indicating the part selected is output.