Avatar Facial Expression Control Using Semantic Masks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine vision systems face challenges in handling dynamic 3D objects, particularly in generating novel views of human faces with varying expressions, due to the complexity of manual annotation and computational requirements, which can be time-consuming and resource-intensive.

Innovation Solution

A method involving automatic detection of facial landmarks and action units (AUs) from a single data source, followed by generating semantic masks and building a facial hyperspace using a neural radiance field, allows for the generation of a synthetic image with modifiable expressions, reducing the need for manual annotation and optimizing processing time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual annotation is used for facial expressions, then control precision of avatar expressions is improved, but time consumption and resource requirements increase significantly

Engineering Contradiction:
Improvecontrol precisionVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces manual annotation (mechanical human operation) with an automated machine vision system that uses neural networks to detect facial landmarks and generate semantic masks, thereby eliminating time-consuming manual work while maintaining expression control precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the avatar to automatically control its own facial expressions through the neural network-based detection and generation pipeline, without requiring external manual annotation for each expression state

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If dynamic 3D objects are rendered with multiple views, then representation accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improverepresentation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-processing input images to detect facial landmarks and generate semantic masks before rendering, which organizes and structures the data in advance, reducing computational complexity during the actual rendering process while maintaining high representation accuracy

Inventive Principle:
Principle #10Preliminary action

3Productivity

If neural networks are used for automatic detection, then productivity is improved, but measurement precision of facial features may be compromised

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidfacial feature detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces semantic masks as an intermediary between the neural network detection and the final avatar rendering. These masks act as a bridge that refines and enhances the detected facial features, ensuring high measurement precision while maintaining the productivity benefits of automated neural network processing

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12406423B2Avatar control
Publication Date: 2025.09.02 FUJITSU LTD
  • US12406423B2 patent drawing
  • US12406423B2 patent drawing
  • US12406423B2 patent drawing

AI summary

In an example, a method may include obtaining, from a data source, first data including multiple frames each including a human face. The method may include automatically detecting, in each of the multiple frames, one or more facial landmarks and one or more action units (AUs) associated with the human face. The method may also include automatically generating one or more semantic masks based at least on the one or more facial landmarks, the one or more semantic masks individually corresponding to the human face. The method may further include obtaining a facial hyperspace using at least the first data, the one or more AUs, and the semantic masks. The method may also include generating a synthetic image of the human face using a first frame of the multiple frames and one or more AU intensities individually associated with the one or more AUs.