Phoneme to Lip Sync Animation Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3D animation technologies face challenges in creating realistic and smooth lip sync animations for characters, as existing phoneme-targeting methods often result in 'choppy' or 'robotic' movements due to the inability to accurately detect and interpret the complex relationships between phonemes and mouth movements.

Innovation Solution

A system and method that converts phoneme transcription data into animation data for 16 independent animation parameters, using a configuration file and algorithmic transformations to produce more realistic and aesthetically pleasing lip sync animations by modifying the timing and interpolation of mouth movements based on phoneme data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If 1-to-1 phoneme targeting is used to control mouth movements, then the lip sync animation can be automatically generated from phoneme data, but the resulting animation appears choppy and robotic

Engineering Contradiction:
Improveautomatic lip sync generationVSAvoidsmoothness of mouth movements
Core Design Contradiction:
Extent of automationVSManufacturing precision

Solution Approach 1:

The patent segments the phoneme sequence into syllable groups and further divides mouth movements into multiple animation parameters (16 blendshape parameters). Instead of mapping each phoneme to a single mouth pose, the system segments the control into temporal groups (syllables) and spatial components (multiple mouth parameters), allowing smoother transitions between states.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts mouth movements by interpolating between keyframes derived from syllable boundaries rather than rigidly following each phoneme. The animation parameters are dynamically modified based on phoneme duration, stress, and position within syllables, creating more natural and less mechanical mouth movements.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If manual phoneme acquisition is performed by animators, then accurate phoneme detection can be achieved, but the process becomes extremely time-consuming

Engineering Contradiction:
Improveaccuracy of phoneme detectionVSAvoidtime consumption for phoneme acquisition
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces an intermediary processing layer between automatic phoneme detection and animation generation. Instead of directly using raw phoneme timestamps, the system processes phoneme data through syllable segmentation, duration analysis, and stress detection algorithms. This intermediary layer enhances the quality of automatic phoneme acquisition by adding contextual information that improves animation accuracy without requiring manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces the manual mechanical process of animator-based phoneme labeling with an automated computational approach. Algorithms automatically detect phoneme boundaries, calculate durations, identify stress patterns, and generate animation keyframes, substituting human labor with automated signal processing and pattern recognition techniques.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of manufacture

If existing phoneme transcription data is directly converted to animation keyframes, then the conversion process is simple, but the resulting lip sync animation lacks naturalness and smoothness

Engineering Contradiction:
Improvesimplicity of conversion processVSAvoidnaturalness of lip sync animation
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The system performs preliminary processing of phoneme data before animation generation. It pre-calculates syllable boundaries, determines phoneme durations, identifies stress patterns, and prepares interpolated keyframe values in advance. This preliminary action ensures that when the actual animation is generated, the data is already optimized for natural-looking movements, maintaining simplicity while improving quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the animation by modifying multiple parameters simultaneously: adjusting keyframe timing based on phoneme duration, interpolating mouth pose values across 16 parameters, modifying transition speeds, and applying stress-based variations. These parameter changes collectively enhance the naturalness of lip sync while building upon the simple phoneme-to-animation conversion framework.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11024071B2Method of converting phoneme transcription data into lip sync animation data for 3D animation software
Publication Date: 2021.06.01 ESPIRITU TECH LLC
  • US11024071B2 patent drawing
  • US11024071B2 patent drawing
  • US11024071B2 patent drawing

AI summary

Described is a system, method, and computer program product that substantially advances the art of animating Lip Sync in 3D computer animated characters by automatically producing data from a Phoneme Transcription of a dialog audio file, which data results in Lip Sync animation that is more realistic, smooth, and aesthetically pleasing than that produced by current Phoneme-Target Lip Sync systems. This Invention works by converting a Phoneme Transcription of a recorded dialog audio file into KeyFrame Data which dynamically controls 16 independent animation Parameters, each associated with a different part of the animated character's mouth, then algorithmically modifying that data such that it conforms to the previously unknown complex, subtle and context-specific relationships between audible phonemes and visible mouth movements.