Speech-Waveform Emotion Recognition for Expressive Avatar Faces

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing communication systems using avatars struggle to express rich emotions due to minimal changes in user facial expressions, limiting the conveyance of nuanced information.

Innovation Solution

An information processing device that includes an emotion recognition unit to analyze speech waveforms, a facial expression output unit to generate corresponding expressions, and an avatar composition unit to control the avatar's emotions, enhancing emotional expression through speech recognition and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If motion capturing is used to generate avatar facial expressions, then the avatar can replicate user expressions, but the avatar cannot express rich emotions due to minimal changes in user facial expressions

Engineering Contradiction:
Improveemotional expression capabilityVSAvoidemotion information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent introduces speech waveform analysis as an intermediary to bridge the gap between limited facial expression changes and rich emotional expression. The speech waveform serves as a mediator that contains emotional information not visible in facial expressions, allowing the avatar to convey emotions that would otherwise be lost

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical motion capturing system with an acoustic analysis system. Instead of relying on physical facial movement capture, the system uses speech waveform analysis to detect emotions, substituting a mechanical measurement approach with an acoustic field-based approach that can detect subtle emotional states

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If speech waveform analysis is added to detect emotions, then the avatar can express rich emotions, but the system complexity increases

Engineering Contradiction:
Improveemotional expression capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent makes the speech processing system multi-functional by using it for both text-to-speech conversion and emotion recognition. The same speech waveform data is analyzed for emotional content, eliminating the need for separate emotion sensing hardware and reducing overall system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the emotion recognition function with the existing speech processing pipeline. By combining emotion detection with text-to-speech processing, the system achieves rich emotional expression without adding separate complex subsystems, as both functions share the same input data stream

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12423897B2Information processing device, information processing method, and program
Publication Date: 2025.09.23 SONY GROUP CORP
  • US12423897B2 patent drawing
  • US12423897B2 patent drawing
  • US12423897B2 patent drawing

AI summary

An information processing device includes an emotion recognition unit, a facial expression output unit, and an avatar composition unit. The emotion recognition unit recognizes an emotion on the basis of a speech waveform. The facial expression output unit outputs a facial expression corresponding to the emotion. The avatar composition unit composes an avatar showing the output facial expression.