Voice Emotion Estimation Using Linguistic-Aware Model Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotion estimation technologies, such as those described in Japanese Unexamined Patent Application Publication No. 2019-28910, do not adequately utilize machine learning to estimate customer emotions during business talks, limiting their effectiveness in improving customer service and support quality.
Innovation Solution
An emotion estimation method that employs a first estimation model based on linguistic information and a second model based on paralinguistic and non-linguistic information, switching between models depending on the presence of linguistic content in voice data to optimize emotion estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single estimation model is used for all voice data, then the system complexity is reduced, but the emotion estimation accuracy deteriorates when linguistic information is not available
Solution Approach 1:
The estimation model is segmented into two distinct models: a first estimation model that processes linguistic information and a second estimation model that processes paralinguistic and non-linguistic information. This segmentation allows each model to be optimized for specific data types, improving overall accuracy without requiring a single overly complex universal model.
Solution Approach 2:
The system dynamically selects which estimation model to use based on the characteristics of the input voice data. When linguistic information is detected, the first model is activated; otherwise, the second model is used. This dynamic adaptation ensures optimal accuracy for each data type while maintaining manageable system complexity through conditional logic.
2Measurement precision
If linguistic information is always processed, then the emotion estimation can utilize semantic content, but the system cannot accurately estimate emotions when linguistic information is absent or unreliable
Solution Approach 1:
The system applies different processing qualities to different parts of the voice data. When linguistic information is present and reliable, it is processed using the first estimation model for semantic emotion analysis. When linguistic information is absent or unreliable, the system switches to processing paralinguistic and non-linguistic features using the second model, ensuring appropriate quality of analysis for each data condition.
Solution Approach 2:
The system changes the processing parameters by switching between two different estimation models based on the availability and quality of linguistic information. This parameter change allows the system to adapt its analysis approach according to the data characteristics, maintaining high accuracy across different voice data scenarios.
3Adaptability or versatility
If only paralinguistic and non-linguistic information is used, then the system can estimate emotion without linguistic content, but the estimation accuracy deteriorates when linguistic information is available
Solution Approach 1:
The estimation functionality is segmented into two specialized models: one for linguistic information processing and one for paralinguistic and non-linguistic information processing. This segmentation enables the system to leverage the appropriate model based on data availability, achieving high accuracy in both linguistic and non-linguistic contexts.
Solution Approach 2:
The system incorporates feedback about the presence and quality of linguistic information to determine which estimation model to use. This feedback mechanism ensures that the most accurate model is selected for each input, maximizing overall estimation accuracy while maintaining versatility across different data types.
Data Source
AI summary
An emotion estimation method that is executed by an information processing device includes: acquiring voice data; determining whether the voice data includes linguistic information; and estimating an emotion corresponding to the voice data by inputting the voice data to a first estimation model, when the voice data includes the linguistic information, or estimating the emotion corresponding to the voice data by inputting the voice data to a second estimation model, when the voice data does not include the linguistic information. The first estimation model estimates the emotion based on the linguistic information. The second estimation model estimates the emotion without being based on the linguistic information.


