3D-2D Face Expression Recognition via Depth and Color Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Face expression recognition accuracy declines due to varying face postures and light conditions, as existing methods rely on two-dimensional image features which are prone to errors in such scenarios.

Innovation Solution

A method and device that utilize a combination of three-dimensional and two-dimensional image processing, including feature point determination, rotation, transformation, alignment, contrast stretching, and normalization, to enhance the accuracy of face expression recognition by inputting depth and color information into a neural network for classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If two-dimensional image feature extraction is used for expression recognition, then the method is simple and fast, but the recognition accuracy declines under varying face postures and light conditions

Engineering Contradiction:
Improveexpression recognition accuracyVSAvoidadaptability to different postures and light conditions
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transitions from two-dimensional image processing to three-dimensional depth image processing. By capturing facial expressions in 3D space using depth cameras, the system obtains accurate geometric information about facial muscle movements that remains consistent across different viewing angles and lighting conditions, thereby resolving the adaptability issue while maintaining recognition accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent combines multiple types of image data including depth information, color information, and infrared information to create a composite representation of facial expressions. This multi-modal fusion approach leverages the complementary strengths of different sensing modalities to achieve robust expression recognition that is insensitive to posture and lighting variations.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If three-dimensional and two-dimensional image processing combination is used, then the expression recognition accuracy is improved, but the device complexity increases

Engineering Contradiction:
Improveexpression recognition accuracyVSAvoidimage processing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the complex processing task into separate modules: one module processes depth images to extract geometric features, another processes color images to extract texture features, and a final module fuses these features for classification. This modular segmentation reduces the complexity of any single processing component while achieving high overall accuracy through their combination.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11023715B2Method and apparatus for expression recognition
Publication Date: 2021.06.01 ARCSOFT CORP LTD
  • US11023715B2 patent drawing
  • US11023715B2 patent drawing
  • US11023715B2 patent drawing

AI summary

The present disclosure provides a method and apparatus for expression recognition, which is applied to the field of image processing. The method includes acquiring a three-dimensional image of a target face and a two-dimensional image of the target face, where the three-dimensional image includes first depth information of the target face and first color information of the target face, and the two-dimensional image includes second color information of the target face. A first neural network classifies an expression of the target face according to the first depth information, the first color information, the second color information, and a first parameter to the target face. The first parameter includes at least one facial expression category and first parameter data for identifying an expression category of the target facial. The disclosed method and device can accurately recognize facial expressions under different facial positions and different illumination conditions.