Autofocus Eye Selection via Depth Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current autofocus systems in cameras, particularly when focusing on humans or animals, face challenges in accurately determining the eye or pupil for sharp focus, especially when the subject is not directly facing the camera, leading to inconsistent focus results.
Innovation Solution
An image processing apparatus and method that utilizes a trained model to estimate the distance between eye and nose features in an image, selecting the eye closer to the camera for focus adjustment, thereby improving autofocus accuracy by leveraging deep neural networks for feature detection and distance estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional autofocus detection methods are used, then the system is simple to implement, but the accuracy of detecting the correct eye for focus is poor when the subject is not directly facing the camera
Solution Approach 1:
The patent transitions from 2D image plane detection to 3D spatial reasoning by estimating depth distances from the camera to each eye. The neural network model predicts distance values for multiple eye candidates, enabling the system to select the closest eye based on inferred depth information rather than relying solely on 2D image coordinates and facial orientation.
Solution Approach 2:
The patent replaces traditional mechanical/optical autofocus detection mechanisms with an AI-based neural network approach. Instead of using phase detection sensors or contrast-based methods that struggle with angled subjects, the system uses deep learning models trained on labeled data to directly predict eye positions and distances, substituting physical detection methods with computational intelligence.
2Measurement precision
If deep neural networks are used for feature detection, then the measurement precision of eye position and distance is improved, but the computational complexity and processing time increase
Solution Approach 1:
The patent applies preliminary action by pre-training neural network models on large datasets of labeled images before deployment. The models are trained offline to learn the complex mappings between image features and eye position/distance relationships, so that during actual autofocus operation, the pre-trained networks can quickly infer distances without requiring complex real-time computations.
Solution Approach 2:
The patent uses copying by creating synthetic training data through image processing and labeling. Ground truth distance maps are generated by processing labeled images through geometric transformations, creating copies of real scenarios with known answers. This allows the neural network to be trained on extensive synthetic data that mirrors real-world conditions without requiring equally extensive real-world annotated datasets.
Data Source
AI summary
An image processing apparatus is provided. First position estimation is performed to detect one or more parts of a first type in an image. A distance between each of the one or more parts of the first type and a part of a second type in the image is estimated. From among the one or more parts of the first type detected, one part of the first type is selected based on the distance estimated for each of the parts of the first type. Information indicating the part selected is output.


