Emotion Estimation Using User-Specific Feature Extraction Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotion estimation techniques heavily rely on the characteristics of a user's way of speaking, leading to inaccurate extraction of emotional fluctuations, as they often fail to account for variations beyond normal speaking patterns.
Innovation Solution
An emotion estimation apparatus that acquires and processes medium data from users, extracts features, and calculates emotion features based on a user-specific emotion feature extraction model, allowing for precise emotion estimation by considering changes from normal speaking patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If features of the user's way of speaking are extracted using a speaker recognition neural network model trained with universal speech data, then the emotion estimation can utilize characteristics of the user's speaking style, but the estimation excessively depends on normal speaking patterns and fails to properly extract emotional fluctuations
Solution Approach 1:
The patent segments the feature extraction process into two distinct components: (1) a speaker recognition neural network model that extracts features of the user's way of speaking from universal speech data, and (2) a separate emotion feature extraction model that is specifically trained to detect emotional fluctuations. This segmentation allows each model to specialize in its respective function, resolving the contradiction between utilizing speaking style characteristics and detecting emotional variations.
Solution Approach 2:
The patent introduces an intermediary emotion feature extraction model that acts as a mediator between the speaker recognition model and the final emotion estimation. This intermediary model takes the features extracted by the speaker recognition model and transforms them into emotion-specific features, enabling the system to leverage speaking style characteristics while maintaining sensitivity to emotional fluctuations.
2Quantity of substance
If a speaker recognition neural network model trained with large amount of universal speech data is used to extract features of the user's way of speaking, then comprehensive speaking characteristics can be captured, but the model cannot distinguish emotional variations from normal speaking patterns
Solution Approach 1:
The patent divides the processing pipeline into two specialized models: the speaker recognition model processes large amounts of universal speech data to capture comprehensive speaking characteristics, while the separate emotion feature extraction model focuses specifically on detecting emotional fluctuations. This segmentation allows each model to optimize for its specific task without interference from the other.
Solution Approach 2:
The emotion feature extraction model is trained in a self-service manner using speech data from the target user, allowing it to adapt to the user's specific characteristics while maintaining the ability to detect emotional variations. This self-training process enables the model to distinguish emotional fluctuations from normal speaking patterns specific to that user.
Data Source
AI summary
According to one embodiment, an emotion estimation apparatus includes processing circuitry. The processing circuitry is configured to acquire medium data of a target user in which a medium of the target user is recorded, extract medium features from the medium data of the target user, calculate an emotion feature based on the extracted medium features and an emotion feature extraction model denoting a reference for the medium features of the target user, and estimate an emotion of the target user based on the calculated emotion feature.


