Speech-Waveform Emotion Recognition for Expressive Avatar Faces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication systems using avatars struggle to express rich emotions due to minimal changes in user facial expressions, limiting the conveyance of nuanced information.
Innovation Solution
An information processing device that includes an emotion recognition unit to analyze speech waveforms, a facial expression output unit to generate corresponding expressions, and an avatar composition unit to control the avatar's emotions, enhancing emotional expression through speech recognition and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If motion capturing is used to generate avatar facial expressions, then the avatar can replicate user expressions, but the avatar cannot express rich emotions due to minimal changes in user facial expressions
Solution Approach 1:
The patent introduces speech waveform analysis as an intermediary to bridge the gap between limited facial expression changes and rich emotional expression. The speech waveform serves as a mediator that contains emotional information not visible in facial expressions, allowing the avatar to convey emotions that would otherwise be lost
Solution Approach 2:
The patent replaces the mechanical motion capturing system with an acoustic analysis system. Instead of relying on physical facial movement capture, the system uses speech waveform analysis to detect emotions, substituting a mechanical measurement approach with an acoustic field-based approach that can detect subtle emotional states
2Adaptability or versatility
If speech waveform analysis is added to detect emotions, then the avatar can express rich emotions, but the system complexity increases
Solution Approach 1:
The patent makes the speech processing system multi-functional by using it for both text-to-speech conversion and emotion recognition. The same speech waveform data is analyzed for emotional content, eliminating the need for separate emotion sensing hardware and reducing overall system complexity
Solution Approach 2:
The patent merges the emotion recognition function with the existing speech processing pipeline. By combining emotion detection with text-to-speech processing, the system achieves rich emotional expression without adding separate complex subsystems, as both functions share the same input data stream
Data Source
AI summary
An information processing device includes an emotion recognition unit, a facial expression output unit, and an avatar composition unit. The emotion recognition unit recognizes an emotion on the basis of a speech waveform. The facial expression output unit outputs a facial expression corresponding to the emotion. The avatar composition unit composes an avatar showing the output facial expression.


