Computer Generated Head Animation with Expression Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer-generated talking heads lack realism in their animation, particularly in mimicking human speech and expression, failing to convincingly convey emotions and speaking styles.
Innovation Solution
A method and system that animate a computer-generated head by converting text or speech into acoustic units, using statistical models to generate image vectors that define facial expressions and movements, allowing for synchronized audio and video output, with adaptive capabilities to mimic various emotions and speaking styles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If statistical models with expression-dependent weights are used to generate image vectors, then the realism and naturalness of facial expressions are improved, but the device complexity and computational requirements increase
Solution Approach 1:
The patent segments the statistical model into multiple clusters, where each cluster represents a specific expression category (e.g., happy, sad, angry). This segmentation allows the system to manage complexity by dividing the overall expression generation task into smaller, more manageable clusters, each with its own parameters and weightings.
Solution Approach 2:
The patent performs preliminary action by pre-computing and storing expression-dependent weightings for each cluster before actual animation generation. During runtime, the system only needs to retrieve and combine these pre-computed weightings according to the detected expression, rather than computing everything from scratch, thus reducing real-time computational complexity.
2Adaptability or versatility
If multiple clusters with expression-dependent weightings are implemented, then the adaptability to different emotions and speaking styles is improved, but the quantity of data and model parameters increases
Solution Approach 1:
The patent implements universality by designing a core statistical model structure that can handle multiple expression types through the cluster framework. Each cluster shares the same underlying model architecture and parameter types, allowing the system to adapt to different expressions (happy, sad, angry, etc.) and speaking styles without requiring completely separate models for each case.
Solution Approach 2:
The patent uses parameter changes by adjusting the weightings of model parameters based on the detected expression category. Instead of changing the entire model structure, the system modifies specific parameter weightings within each cluster to reflect different expressions, enabling versatile expression adaptation while keeping the base parameter set manageable.
3Manufacturing precision
If synchronized audio and video output with lip synchronization is implemented, then the realism of speech animation is improved, but the complexity of coordination between audio and video processing increases
Solution Approach 1:
The patent implements feedback by using the audio input as a reference to drive the video generation process. The acoustic unit detection and expression recognition from audio create a feedback loop that continuously guides the image vector generation, ensuring that lip movements remain synchronized with speech content. The system constantly adjusts video output based on real-time audio analysis.
Solution Approach 2:
The patent introduces an intermediary process - the statistical model that converts acoustic units into image vectors. This intermediary transformation layer bridges the audio and video domains, translating audio features (acoustic units) into visual features (facial expressions and lip movements) through learned statistical relationships, thereby simplifying the coordination between audio and video processing.
Data Source
AI summary
A method of animating a computer generation of a head, the head having a mouth which moves in accordance with speech to be output by the head,said method comprising:providing an input related to the speech which is to be output by the movement of the lips;dividing said input into a sequence of acoustic units;selecting expression characteristics for the inputted text;converting said sequence of acoustic units to a sequence of image vectors using a statistical model, wherein said model has a plurality of model parameters describing probability distributions which relate an acoustic unit to an image vector, said image vector comprising a plurality of parameters which define a face of said head; andoutputting said sequence of image vectors as video such that the mouth of said head moves to mime the speech associated with the input text with the selected expression,wherein a parameter of a predetermined type of each probability distribution in said selected expression is expressed as a weighted sum of parameters of the same type, and wherein the weighting used is expression dependent, such that converting said sequence of acoustic units to a sequence of image vectors comprises retrieving the expression dependent weights for said selected expression, wherein the parameters are provided in clusters, and each cluster comprises at least one sub-cluster, wherein said expression dependent weights are retrieved for each cluster such that there is one weight per sub-cluster.


