3D Face Model Generation Using Selective Expression Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating highly-accurate three-dimensional (3D) models for photorealistic facial animation are burdensome, requiring extensive photography setups and subject participation in making numerous facial expressions, which can be stressful and impractical.

Innovation Solution

An image processing device and method that efficiently generate highly-accurate 3D face models by selectively shooting facial expressions with unique characteristics and analyzing the shooting state in real-time to reduce errors and improve processing stability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multiple cameras are placed around the face to shoot various facial expressions, then the accuracy of 3D model generation is improved, but the device complexity and burden on the photographer increase significantly

Engineering Contradiction:
Improve3D model accuracyVSAvoidphotography setup complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses a single camera to capture facial expressions and then creates multiple virtual camera views through image processing and 3D reconstruction algorithms. This copying approach replicates the effect of multiple physical cameras without the associated complexity, burden, and cost of setting up numerous actual cameras around the subject's face.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If the subject makes many different facial expressions to generate target shapes, then the photorealistic quality of the animation is improved, but the mental and physical stress on the subject increases

Engineering Contradiction:
Improvephotorealistic qualityVSAvoidsubject stress
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent performs comprehensive 3D model generation and expression analysis in advance during the photography session. The system captures a wide range of facial expressions and pre-processes this data to create a robust 3D model and expression library. This preliminary action reduces the need for extensive additional expression shooting later, thereby reducing subject stress while maintaining photorealistic quality.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If feature point detection is used to generate target shapes, then the device complexity is reduced, but the measurement precision and reliability of the 3D model decrease

Engineering Contradiction:
Improvephotography setup simplicityVSAvoidtarget shape accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical feature point detection methods with advanced image processing and 3D reconstruction algorithms. Instead of relying solely on detecting specific facial landmarks, the system uses comprehensive image analysis, depth information processing, and geometric reconstruction to generate accurate 3D models. This substitution maintains simplicity in the photography setup while significantly improving measurement precision and model reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12266042B2Image processing device and image processing method
Publication Date: 2025.04.01 SONY GROUP CORP
  • US12266042B2 patent drawing
  • US12266042B2 patent drawing
  • US12266042B2 patent drawing

AI summary

Provided are a device and method that enable highly accurate and efficient three-dimensional model generation processing. The device includes: a facial feature information detection unit that analyzes a facial image of a subject shot by an image capturing unit and detects facial feature information; an input data selection unit that selects, from a plurality of facial images shot by the image capturing unit and a plurality of pieces of facial feature information corresponding to the plurality of facial images, a set of a facial image and feature information optimal for generating a 3D model; and a facial expression 3D model generation unit that generates a 3D model using the facial image and the feature information selected by the input data selection unit. As the data optimal for generating a 3D model, the input data selection unit selects, for example, a facial image having feature information with a large change from standard data constituted by an expressionless 3D model and high reliability, as well as feature information.