Emotion Estimation Apparatus Speech State Model Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing emotion estimation technologies face challenges in accurately estimating emotions from facial images, particularly when the target individual is speaking or not, as they fail to differentiate between speech and non-speech states effectively, leading to inconsistent emotion recognition results.

Innovation Solution

An emotion estimation apparatus that employs a processor to execute a speech determination process and an emotion estimation process using facial images, switching between two distinct emotion recognition models based on speech determination results, one focusing on the mouth area when the individual is not speaking and another excluding the mouth area when they are speaking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single emotion recognition model is used for both speech and non-speech states, then the device complexity is reduced, but the emotion recognition accuracy deteriorates due to inconsistent results across different speech states

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidmodel structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the emotion recognition task into two separate models: a first emotion recognition model for non-speech states and a second emotion recognition model for speech states. Each model is specialized to handle specific speech conditions, thereby improving recognition accuracy without requiring a single complex universal model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically switches between different emotion recognition models based on the detected speech state. The control unit determines whether the target is speaking or not and selects the appropriate model, allowing the system to adapt its structure to match the current operational conditions for optimal accuracy.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If the mouth area is always included in emotion recognition, then more facial information is available for non-speech states, but recognition accuracy deteriorates when the individual is speaking due to mouth movements

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidadaptability to speech states
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies different processing strategies to different facial regions based on speech state. For non-speech states, the mouth area is included to capture full facial expressions. For speech states, the mouth area is excluded or given reduced weight to avoid interference from mouth movements, while other facial regions maintain their processing.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the input parameters and processing weights of the emotion recognition model based on speech state detection. When speech is detected, the model adjusts by reducing or eliminating the contribution of mouth region features, thereby adapting to the changing facial dynamics during speech while maintaining accuracy.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If emotion estimation is performed without speech determination, then the processing speed is improved, but the reliability of emotion estimation deteriorates due to inconsistent results

Engineering Contradiction:
Improveemotion estimation consistencyVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs speech determination as a preliminary step before executing emotion estimation. By detecting the speech state in advance, the system can select the appropriate emotion recognition model and adjust processing parameters beforehand, ensuring consistent and reliable results while maintaining efficient processing through pre-planned execution paths.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10255487B2Emotion estimation apparatus using facial images of target individual, emotion estimation method, and non-transitory computer readable medium
Publication Date: 2019.04.09 CASIO COMPUTER CO LTD
  • US10255487B2 patent drawing
  • US10255487B2 patent drawing
  • US10255487B2 patent drawing

AI summary

A speech determiner determines whether or not a target individual is speaking when facial images of the target individual are captured. An emotion estimator estimates the emotion of the target individual using the facial images of the target individual, on the basis of the determination results of the speech determiner.