Singing Viseme Animation Curves With Melodic Accent Modulation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for transforming vocal performances into visual performances, particularly for singing, are challenged by the ill-suited lip-sync timing and mouth shapes designed for speech visualization, failing to accurately depict the expressive nature of singing due to the different roles of vowels and consonants in conveying melody and rhythm.
Innovation Solution
A system and method that modulate animation curves using Melodic-accent (Ma) and Pitch-sensitivity (Ps) parameters to dynamically represent singing styles, considering the physiological and stylistic aspects of vocal performance, with a four-dimensional (4D) space encompassing Jaw (Ja) and Lip (Li) parameterization to decouple and layer spatio-temporal contributions of vowels and consonants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If speech-based lip-sync timing and mouth shapes are used for singing visualization, then the system can be simple and based on existing speech models, but the accuracy of depicting singing expression and melody is poor
Solution Approach 1:
The patent segments the animation control into multiple independent dimensions: jaw movement (Ja), lip shape (Li), melodic accent (Ma), and pitch sensitivity (Ps). This segmentation allows each parameter to be controlled independently to accurately represent different aspects of singing, resolving the contradiction by creating a structured multi-parameter system that maintains organization while achieving high precision.
Solution Approach 2:
The patent extends the traditional 2D viseme space by adding two new dimensions (Ma and Ps), creating a 4D parameter space. This dimensional expansion enables precise control over singing-specific characteristics like melody emphasis and pitch variation, achieving accurate singing visualization while maintaining systematic control through the extended parameter framework.
2Adaptability or versatility
If speech-based viseme models are used for singing animation, then the ease of implementation is high, but the adaptability to singing styles and expressions is poor
Solution Approach 1:
The patent makes the animation parameters dynamic by allowing continuous adjustment of Ja, Li, Ma, and Ps values based on the singing input. This dynamic control enables the system to adapt to different singing styles, emotions, and expressions in real-time, achieving high versatility while maintaining systematic control through the defined parameter relationships.
Solution Approach 2:
The patent creates a universal animation framework that can handle multiple singing styles and expressions through a single set of four parameters. This multi-functional system can represent various singing genres, emotional states, and stylistic choices by adjusting the parameter combinations, achieving broad adaptability without requiring separate models for each case.
3Loss of information
If traditional speech animation curves are used, then the simplicity of the animation system is maintained, but the ability to portray melodic accent and pitch sensitivity is insufficient
Solution Approach 1:
The patent changes the parameter structure from traditional speech animation by introducing Ma (melodic accent) and Ps (pitch sensitivity) parameters alongside the traditional Ja and Li parameters. This parameter transformation enables the animation curves to encode and convey melodic and pitch information, achieving complete information representation while maintaining systematic control through the defined parameter relationships.
Data Source
AI summary
A system and method of modulating animation curves based on audio input. The method including: identifying phonetic features for a plurality of visemes in the audio input; determining viseme animation curves based on parameters representing a spatial appearance of the plurality of visemes; modulating the viseme animation curves based on melodic accent, pitch sensitivity, or both, based on the phonetic features; and outputting the modulated animation curves.


