Real-time Speech Animation via Hierarchical Snippet Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech animation techniques struggle with real-time generation of realistic and varied speech animations, particularly in accommodating co-articulation effects and emotional expressions, leading to robotic and repetitive animations.

Innovation Solution

A hierarchical search algorithm is used to retrieve and combine Animation Snippets from a Lookup Table, allowing for the creation of smooth and contextually accurate speech animations. Additionally, model visemes are blended with data-driven animation to enhance realism, and an Output Weighting Function is employed to integrate emotional expressions naturally.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If procedural speech animation with rules or look-up tables is used, then real-time generation is achieved, but animation quality becomes robotic and repetitive

Engineering Contradiction:
Improvereal-time generation speedVSAvoidanimation quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments speech animation into phoneme-level units, where each phoneme can be independently animated with high-quality motion capture data. This segmentation allows real-time assembly of pre-captured high-quality phoneme animations to form complete speech animations, resolving the contradiction between real-time generation and animation quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary motion capture and animation generation for individual phonemes offline, storing these pre-generated animation snippets in a library. During real-time speech animation, these pre-prepared phoneme animations are retrieved and concatenated, eliminating the need for real-time complex computation while maintaining high quality.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If data-driven statistical methods are used, then animation quality improves through stitching facial animation data, but control is taken away from animators and quality is limited by available data

Engineering Contradiction:
Improveanimation qualityVSAvoidanimator control flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a hybrid system that dynamically combines data-driven statistical animation for phoneme transitions with animator-controlled parameters for emotional expression and style. This allows the system to adapt between automated high-quality phoneme animation and animator-directed expressive control based on real-time needs.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces an intermediary control layer that translates animator inputs into adjustments of pre-computed animation data. This intermediary layer preserves animator control while leveraging the quality of data-driven methods, allowing animators to guide emotional expression without directly manipulating raw animation data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If performance-capture based speech animation is used, then realistic facial dynamics are achieved, but the system is complex and real-time playback is difficult

Engineering Contradiction:
Improvefacial dynamics realismVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential phoneme-specific animation data from complex performance capture sessions, separating the core speech-related facial movements from emotional and contextual variations. This extraction creates a simplified phoneme animation library that maintains realism while reducing system complexity for real-time application.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms complex performance capture data into a standardized parameter representation where each phoneme is defined by a small set of key animation parameters. This parameterization simplifies the data structure and enables efficient real-time retrieval and combination while preserving the realism of original performance capture.

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If prior speech animation techniques are used, then speech animation is achieved, but emotional expression animation cannot be adequately accommodated

Engineering Contradiction:
Improvespeech animation qualityVSAvoidemotional expression capability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal phoneme animation framework that can simultaneously accommodate speech animation and emotional expression animation. The system treats emotional expressions as additional layers that can be composited with phoneme-based speech animation, allowing a single system to handle both functions without compromising either.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges speech animation data and emotional expression animation data into a unified animation output. By combining phoneme-level speech movements with emotion-level facial expressions in a coordinated manner, the system achieves both accurate speech articulation and natural emotional expression simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12315054B2Real-time generation of speech animation
Publication Date: 2025.05.27 SOUL MACHINES LTD
  • US12315054B2 patent drawing
  • US12315054B2 patent drawing
  • US12315054B2 patent drawing

AI summary

To realistically animate a String (such as a sentence) a hierarchical search algorithm is provided to search for stored examples (Animation Snippets) of sub-strings of the String, in decreasing order of sub-string length, and concatenate retrieved sub-strings to complete the String of speech animation. In one embodiment, real-time generation of speech animation uses model visemes to predict the animation sequences at onsets of visemes and a look-up table based (data-driven) algorithm to predict the dynamics at transitions of visemes. Specifically posed Model Visemes may be blended with speech animation generated using another method at corresponding time points in the animation when the visemes are to be expressed. An Output Weighting Function is used to map Speech input and Expression input into Muscle-Based Descriptor weightings.