Voice Emotion Estimation Using Linguistic-Aware Model Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing emotion estimation technologies, such as those described in Japanese Unexamined Patent Application Publication No. 2019-28910, do not adequately utilize machine learning to estimate customer emotions during business talks, limiting their effectiveness in improving customer service and support quality.

Innovation Solution

An emotion estimation method that employs a first estimation model based on linguistic information and a second model based on paralinguistic and non-linguistic information, switching between models depending on the presence of linguistic content in voice data to optimize emotion estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single estimation model is used for all voice data, then the system complexity is reduced, but the emotion estimation accuracy deteriorates when linguistic information is not available

Engineering Contradiction:
Improveemotion estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The estimation model is segmented into two distinct models: a first estimation model that processes linguistic information and a second estimation model that processes paralinguistic and non-linguistic information. This segmentation allows each model to be optimized for specific data types, improving overall accuracy without requiring a single overly complex universal model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects which estimation model to use based on the characteristics of the input voice data. When linguistic information is detected, the first model is activated; otherwise, the second model is used. This dynamic adaptation ensures optimal accuracy for each data type while maintaining manageable system complexity through conditional logic.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If linguistic information is always processed, then the emotion estimation can utilize semantic content, but the system cannot accurately estimate emotions when linguistic information is absent or unreliable

Engineering Contradiction:
Improveemotion estimation accuracyVSAvoidadaptability to different voice data types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system applies different processing qualities to different parts of the voice data. When linguistic information is present and reliable, it is processed using the first estimation model for semantic emotion analysis. When linguistic information is absent or unreliable, the system switches to processing paralinguistic and non-linguistic features using the second model, ensuring appropriate quality of analysis for each data condition.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the processing parameters by switching between two different estimation models based on the availability and quality of linguistic information. This parameter change allows the system to adapt its analysis approach according to the data characteristics, maintaining high accuracy across different voice data scenarios.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If only paralinguistic and non-linguistic information is used, then the system can estimate emotion without linguistic content, but the estimation accuracy deteriorates when linguistic information is available

Engineering Contradiction:
Improveability to estimate emotion without linguistic contentVSAvoidemotion estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The estimation functionality is segmented into two specialized models: one for linguistic information processing and one for paralinguistic and non-linguistic information processing. This segmentation enables the system to leverage the appropriate model based on data availability, achieving high accuracy in both linguistic and non-linguistic contexts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system incorporates feedback about the presence and quality of linguistic information to determine which estimation model to use. This feedback mechanism ensures that the most accurate model is selected for each input, maximizing overall estimation accuracy while maintaining versatility across different data types.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20260080893A1Emotion estimation method, information processing device, and non-transitory storage medium
Publication Date: 2026.03.19 TOYOTA JIDOSHA KK
  • US20260080893A1 patent drawing
  • US20260080893A1 patent drawing
  • US20260080893A1 patent drawing

AI summary

An emotion estimation method that is executed by an information processing device includes: acquiring voice data; determining whether the voice data includes linguistic information; and estimating an emotion corresponding to the voice data by inputting the voice data to a first estimation model, when the voice data includes the linguistic information, or estimating the emotion corresponding to the voice data by inputting the voice data to a second estimation model, when the voice data does not include the linguistic information. The first estimation model estimates the emotion based on the linguistic information. The second estimation model estimates the emotion without being based on the linguistic information.