Neural Mouth Shape Generation for Accurate Lip Motion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies fail to produce accurate and realistic lip motion in digital humans and avatars, particularly for consonants, which is exacerbated when displayed on large screens, highlighting visual irregularities.

Innovation Solution

A neural network model generates visual representations of mouth regions using mouth shape data, specifically viseme coefficients, to create high-quality images or meshes with accurate lip motion, independent of facial landmark data, utilizing image-to-image translation networks and U-Net neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If deepfake technology is used to generate realistic images, then visual realism is improved, but lip motion accuracy deteriorates

Engineering Contradiction:
Improvevisual realismVSAvoidlip motion accuracy
Core Design Contradiction:
Illumination intensityVSManufacturing precision

Solution Approach 1:

The patent segments the face into multiple regions including mouth region, teeth region, and tongue region, each processed independently by dedicated neural networks. This segmentation allows precise control of lip motion while maintaining overall visual realism, resolving the contradiction between realistic appearance and accurate lip motion.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representations including 3D morphable models, facial landmark data, and mesh structures as mediators between the input image and final output. These intermediaries enable precise manipulation of lip geometry and motion while preserving the realistic appearance generated by deepfake technology.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If 3D models are used to improve lip geometry accuracy, then lip motion precision is improved, but visual realism deteriorates due to synthetic appearance

Engineering Contradiction:
Improvelip geometry accuracyVSAvoidvisual realism
Core Design Contradiction:
Manufacturing precisionVSIllumination intensity

Solution Approach 1:

The patent merges multiple techniques including deepfake-generated realistic textures, 3D morphable models for geometric accuracy, and neural network refinement. This combination integrates the visual realism of deepfakes with the geometric precision of 3D models, resolving the contradiction between appearance quality and motion accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite representation combining 2D image data, 3D mesh structures, and neural network features. This composite approach integrates the strengths of different methodologies, producing output that simultaneously achieves visual realism and geometric accuracy.

Inventive Principle:
Principle #40Composite materials

3Area of stationary object

If digital humans are displayed on large screens, then visibility is improved, but visual irregularities are worsened

Engineering Contradiction:
Improvedisplay sizeVSAvoidvisual irregularity
Core Design Contradiction:
Area of stationary objectVSManufacturing precision

Solution Approach 1:

The patent applies local quality enhancement by processing the mouth region with higher resolution and more sophisticated neural networks compared to other facial regions. This localized attention ensures that critical areas like lips maintain high precision even when displayed on large screens where irregularities would be more visible.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12361621B2Creating images, meshes, and talking animations from mouth shape data
Publication Date: 2025.07.15 SAMSUNG ELECTRONICS CO LTD
  • US12361621B2 patent drawing
  • US12361621B2 patent drawing
  • US12361621B2 patent drawing

AI summary

Creating images and animations of lip motion from mouth shape data includes providing, as one or more input features to a neural network model, a vector of a plurality of coefficients. Each vector of the plurality of coefficients corresponds to a different mouth shape. Using the neural network model, a data structure output specifying a visual representation of a mouth including lips having a shape corresponding to the vector is generated.