Virtual Character Nonverbal Movement Generation via Multi-Aspect Speech Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for animating virtual characters are limited in generating nonverbal movements that effectively convey mental states, as they primarily focus on keyword detection and rhythmic features, neglecting comprehensive communication through the body, such as hand movements and upper body leaning.
Innovation Solution
A system that analyzes acoustic, syntactic, semantic, and rhetorical aspects of a virtual character's speech to generate mental state indicators, which are then mapped to nonverbal behaviors like head movements, facial expressions, and gestures, incorporating audio features like agitation and emphasis to create a more nuanced and expressive performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If keyword detection methods are used to generate nonverbal movements, then the system can detect specific words and trigger predefined behaviors, but the system is limited in conveying comprehensive mental states and emotional nuance
Solution Approach 1:
The patent segments the analysis of speech into multiple independent components: acoustic analysis (prosody, pitch, energy), syntactic analysis (sentence structure, clauses), semantic analysis (meaning, metaphors), and rhetorical analysis (emphasis, figures of speech). Each component generates separate indicators that are later integrated to form comprehensive mental state indicators, allowing the system to capture nuanced emotional states without requiring a single overly complex analysis module.
Solution Approach 2:
The patent merges multiple analysis streams (acoustic, syntactic, semantic, rhetorical) and their respective indicators into unified mental state indicators. For example, acoustic indicators about speech energy are combined with semantic indicators about metaphorical content to generate comprehensive mental states that reflect both emotional intensity and cognitive meaning, enabling richer nonverbal expression than any single analysis type could achieve alone.
2Adaptability or versatility
If rhythmic features from audio are used to drive nonverbal movements, then beat gestures can be generated, but comprehensive body communication including hand movements and upper body leaning is neglected
Solution Approach 1:
The system segments nonverbal behavior generation into distinct categories: head movements (driven by prosodic features), facial expressions (driven by emotional indicators), hand gestures (driven by both rhythmic and semantic indicators), and upper body leaning (driven by emphasis indicators). This segmentation allows each body part to be controlled by appropriate features from the multiple analysis streams, achieving comprehensive body communication without requiring a monolithic complex system.
Solution Approach 2:
The patent creates a universal mental state indicator framework that serves multiple functions simultaneously: it drives lip sync movements, head movements, facial expressions, hand gestures, and upper body leaning. The same set of analyzed features (acoustic, syntactic, semantic, rhetorical) generates indicators that are mapped to various nonverbal behaviors, making the system efficiently multi-functional without proportionally increasing complexity.
3Loss of information
If keyword-to-behavior rules are used, then specific gestures can be triggered by detected keywords, but the system cannot convey nuanced emotional states and communicative intent
Solution Approach 1:
Instead of relying on simple keyword detection, the patent segments the meaning extraction process into multiple analytical layers: acoustic analysis captures emotional tone and intensity, syntactic analysis captures sentence-level intent, semantic analysis captures conceptual meaning including metaphors, and rhetorical analysis captures emphasis and stylistic intent. Each layer contributes specific information that is preserved in the resulting mental state indicators, preventing information loss that would occur with simple keyword matching.
Solution Approach 2:
The patent introduces mental state indicators as intermediary representations between the raw speech analysis and the final nonverbal behavior generation. These indicators serve as a rich intermediate layer that preserves nuanced information from all analysis types (acoustic energy, syntactic structure, semantic meaning, rhetorical emphasis) and translates them into comprehensive behavioral instructions, preventing the information loss that would occur with direct keyword-to-behavior mapping.
Data Source
AI summary
Programs for creating a set of behaviors for lip sync movements and nonverbal communication may include analyzing a character's speaking behavior through the use of acoustic, syntactic, semantic, pragmatic, and rhetorical analyses of the utterance. For example, a non-transitory, tangible, computer-readable storage medium may contain a program of instructions that cause a computer system running the program of instructions to: receive a text specifying words to be spoken by a virtual character; extract metaphoric elements, discourse elements, or both from the text; generate one or more mental state indicators based on the metaphoric elements, the discourse elements, or both; map each of the one or more mental state indicators to a behavior that the virtual character should display with nonverbal movements that convey the mental state indicators; and generate a set of instructions for the nonverbal movements based on the behaviors.


