Monocular Camera Defocus for Accurate Face Scale Measurement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing monocular cameras struggle to accurately determine the true scale of a user's face for virtual try-on and augmented reality applications due to image scale ambiguity, as depth information is not readily available, making it difficult to select the correct product size.

Innovation Solution

The method employs a monocular camera to estimate face size by leveraging image defocus, using autofocus or focal stacks to calculate depth, combined with face mesh technology for improved accuracy, without requiring a dense depth map.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a monocular camera is used for face measurement, then device complexity is reduced, but measurement precision deteriorates due to image scale ambiguity

Engineering Contradiction:
Improvecamera systemVSAvoidface scale
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary computational process that uses defocus information as a mediator to infer depth from a monocular image. By analyzing the blur amount at different regions of the face, the system creates a depth map that resolves the scale ambiguity problem, allowing accurate measurement without complex multi-camera systems

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter being analyzed from sharp focus to defocus blur. Instead of requiring a perfectly focused image, the system leverages the blur amount as a useful signal to determine relative depth. This parameter transformation allows a simple monocular camera to provide depth information that would otherwise require complex hardware

Inventive Principle:
Principle #35Parameter changes

2Productivity

If depth information is calculated from sparse feature points, then processing speed is improved, but measurement precision deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoiddepth accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the face into multiple regions with different depth characteristics (foreground features like eyes and mouth versus background features). By calculating defocus separately for each segment and applying appropriate depth estimates, the system achieves both computational efficiency and accurate depth reconstruction across the entire face

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from 2D image coordinates to 3D depth estimation by incorporating the defocus dimension. The blur amount provides an additional dimensional cue that allows the system to infer depth information from a single 2D image, effectively adding a depth dimension without requiring multiple cameras or complex hardware

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Accurately determines the true scale of a user's face using a monocular camera, enabling precise virtual try-on and augmented reality experiences by calculating depth from sparse feature points and applying face mesh technology.

Implementation Method 1

The method employs a monocular camera to estimate face size by leveraging image defocus, using autofocus or focal stacks to calculate depth

Methodology Applied
Scientific EffectImage defocus: Depth of Field

Data Source

PatentUS12462418B2Monocular camera defocus face measuring
Publication Date: 2025.11.04 SNAP INC
  • US12462418B2 patent drawing
  • US12462418B2 patent drawing
  • US12462418B2 patent drawing

AI summary

A device that measures a size of a user's face, referred to as face scaling, using a monocular camera. Depth is calculated from sparse feature points. A face mesh is used to improve the estimation accuracy. A processing pipeline detects face features by applying a face landmark detection algorithm to find the important face feature points such as the eyes, nose, and mouth. The processing pipeline estimates feature points depth using depth obtained through image defocus. The processing pipeline further scales the face using an estimated depth of the face features.