Emotion Estimation Apparatus Speech State Model Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotion estimation technologies face challenges in accurately estimating emotions from facial images, particularly when the target individual is speaking or not, as they fail to differentiate between speech and non-speech states effectively, leading to inconsistent emotion recognition results.
Innovation Solution
An emotion estimation apparatus that employs a processor to execute a speech determination process and an emotion estimation process using facial images, switching between two distinct emotion recognition models based on speech determination results, one focusing on the mouth area when the individual is not speaking and another excluding the mouth area when they are speaking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single emotion recognition model is used for both speech and non-speech states, then the device complexity is reduced, but the emotion recognition accuracy deteriorates due to inconsistent results across different speech states
Solution Approach 1:
The patent divides the emotion recognition task into two separate models: a first emotion recognition model for non-speech states and a second emotion recognition model for speech states. Each model is specialized to handle specific speech conditions, thereby improving recognition accuracy without requiring a single complex universal model.
Solution Approach 2:
The system dynamically switches between different emotion recognition models based on the detected speech state. The control unit determines whether the target is speaking or not and selects the appropriate model, allowing the system to adapt its structure to match the current operational conditions for optimal accuracy.
2Measurement precision
If the mouth area is always included in emotion recognition, then more facial information is available for non-speech states, but recognition accuracy deteriorates when the individual is speaking due to mouth movements
Solution Approach 1:
The patent applies different processing strategies to different facial regions based on speech state. For non-speech states, the mouth area is included to capture full facial expressions. For speech states, the mouth area is excluded or given reduced weight to avoid interference from mouth movements, while other facial regions maintain their processing.
Solution Approach 2:
The system changes the input parameters and processing weights of the emotion recognition model based on speech state detection. When speech is detected, the model adjusts by reducing or eliminating the contribution of mouth region features, thereby adapting to the changing facial dynamics during speech while maintaining accuracy.
3Reliability
If emotion estimation is performed without speech determination, then the processing speed is improved, but the reliability of emotion estimation deteriorates due to inconsistent results
Solution Approach 1:
The system performs speech determination as a preliminary step before executing emotion estimation. By detecting the speech state in advance, the system can select the appropriate emotion recognition model and adjust processing parameters beforehand, ensuring consistent and reliable results while maintaining efficient processing through pre-planned execution paths.
Data Source
AI summary
A speech determiner determines whether or not a target individual is speaking when facial images of the target individual are captured. An emotion estimator estimates the emotion of the target individual using the facial images of the target individual, on the basis of the determination results of the speech determiner.


