Computer Generated Head Animation with Expression Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer-generated talking heads lack realism in their animation, particularly in mimicking human speech and expression, failing to convincingly convey emotions and speaking styles.

Innovation Solution

A method and system that animate a computer-generated head by converting text or speech into acoustic units, using statistical models to generate image vectors that define facial expressions and movements, allowing for synchronized audio and video output, with adaptive capabilities to mimic various emotions and speaking styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If statistical models with expression-dependent weights are used to generate image vectors, then the realism and naturalness of facial expressions are improved, but the device complexity and computational requirements increase

Engineering Contradiction:
Improverealism of facial expressionsVSAvoidcomplexity of statistical model
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the statistical model into multiple clusters, where each cluster represents a specific expression category (e.g., happy, sad, angry). This segmentation allows the system to manage complexity by dividing the overall expression generation task into smaller, more manageable clusters, each with its own parameters and weightings.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-computing and storing expression-dependent weightings for each cluster before actual animation generation. During runtime, the system only needs to retrieve and combine these pre-computed weightings according to the detected expression, rather than computing everything from scratch, thus reducing real-time computational complexity.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multiple clusters with expression-dependent weightings are implemented, then the adaptability to different emotions and speaking styles is improved, but the quantity of data and model parameters increases

Engineering Contradiction:
Improveadaptability to expressionsVSAvoidnumber of model parameters
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent implements universality by designing a core statistical model structure that can handle multiple expression types through the cluster framework. Each cluster shares the same underlying model architecture and parameter types, allowing the system to adapt to different expressions (happy, sad, angry, etc.) and speaking styles without requiring completely separate models for each case.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter changes by adjusting the weightings of model parameters based on the detected expression category. Instead of changing the entire model structure, the system modifies specific parameter weightings within each cluster to reflect different expressions, enabling versatile expression adaptation while keeping the base parameter set manageable.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If synchronized audio and video output with lip synchronization is implemented, then the realism of speech animation is improved, but the complexity of coordination between audio and video processing increases

Engineering Contradiction:
Improvelip synchronization accuracyVSAvoidcomplexity of audio-video coordination
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements feedback by using the audio input as a reference to drive the video generation process. The acoustic unit detection and expression recognition from audio create a feedback loop that continuously guides the image vector generation, ensuring that lip movements remain synchronized with speech content. The system constantly adjusts video output based on real-time audio analysis.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary process - the statistical model that converts acoustic units into image vectors. This intermediary transformation layer bridges the audio and video domains, translating audio features (acoustic units) into visual features (facial expressions and lip movements) through learned statistical relationships, thereby simplifying the coordination between audio and video processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9959657B2Computer generated head
Publication Date: 2018.05.01 KK TOSHIBA
  • US9959657B2 patent drawing
  • US9959657B2 patent drawing
  • US9959657B2 patent drawing

AI summary

A method of animating a computer generation of a head, the head having a mouth which moves in accordance with speech to be output by the head,said method comprising:providing an input related to the speech which is to be output by the movement of the lips;dividing said input into a sequence of acoustic units;selecting expression characteristics for the inputted text;converting said sequence of acoustic units to a sequence of image vectors using a statistical model, wherein said model has a plurality of model parameters describing probability distributions which relate an acoustic unit to an image vector, said image vector comprising a plurality of parameters which define a face of said head; andoutputting said sequence of image vectors as video such that the mouth of said head moves to mime the speech associated with the input text with the selected expression,wherein a parameter of a predetermined type of each probability distribution in said selected expression is expressed as a weighted sum of parameters of the same type, and wherein the weighting used is expression dependent, such that converting said sequence of acoustic units to a sequence of image vectors comprises retrieving the expression dependent weights for said selected expression, wherein the parameters are provided in clusters, and each cluster comprises at least one sub-cluster, wherein said expression dependent weights are retrieved for each cluster such that there is one weight per sub-cluster.