Emotion Estimation Using User-Specific Feature Extraction Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing emotion estimation techniques heavily rely on the characteristics of a user's way of speaking, leading to inaccurate extraction of emotional fluctuations, as they often fail to account for variations beyond normal speaking patterns.

Innovation Solution

An emotion estimation apparatus that acquires and processes medium data from users, extracts features, and calculates emotion features based on a user-specific emotion feature extraction model, allowing for precise emotion estimation by considering changes from normal speaking patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If features of the user's way of speaking are extracted using a speaker recognition neural network model trained with universal speech data, then the emotion estimation can utilize characteristics of the user's speaking style, but the estimation excessively depends on normal speaking patterns and fails to properly extract emotional fluctuations

Engineering Contradiction:
Improveemotion estimation precisionVSAvoidability to detect emotional fluctuations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the feature extraction process into two distinct components: (1) a speaker recognition neural network model that extracts features of the user's way of speaking from universal speech data, and (2) a separate emotion feature extraction model that is specifically trained to detect emotional fluctuations. This segmentation allows each model to specialize in its respective function, resolving the contradiction between utilizing speaking style characteristics and detecting emotional variations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary emotion feature extraction model that acts as a mediator between the speaker recognition model and the final emotion estimation. This intermediary model takes the features extracted by the speaker recognition model and transforms them into emotion-specific features, enabling the system to leverage speaking style characteristics while maintaining sensitivity to emotional fluctuations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If a speaker recognition neural network model trained with large amount of universal speech data is used to extract features of the user's way of speaking, then comprehensive speaking characteristics can be captured, but the model cannot distinguish emotional variations from normal speaking patterns

Engineering Contradiction:
Improveamount of speech data processedVSAvoidemotional fluctuation detection accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent divides the processing pipeline into two specialized models: the speaker recognition model processes large amounts of universal speech data to capture comprehensive speaking characteristics, while the separate emotion feature extraction model focuses specifically on detecting emotional fluctuations. This segmentation allows each model to optimize for its specific task without interference from the other.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The emotion feature extraction model is trained in a self-service manner using speech data from the target user, allowing it to adapt to the user's specific characteristics while maintaining the ability to detect emotional variations. This self-training process enables the model to distinguish emotional fluctuations from normal speaking patterns specific to that user.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240290346A1Emotion estimation apparatus and emotion estimation method
Publication Date: 2024.08.29 KK TOSHIBA
  • US20240290346A1 patent drawing
  • US20240290346A1 patent drawing
  • US20240290346A1 patent drawing

AI summary

According to one embodiment, an emotion estimation apparatus includes processing circuitry. The processing circuitry is configured to acquire medium data of a target user in which a medium of the target user is recorded, extract medium features from the medium data of the target user, calculate an emotion feature based on the extracted medium features and an emotion feature extraction model denoting a reference for the medium features of the target user, and estimate an emotion of the target user based on the calculated emotion feature.