Emotion Recognition via Phoneme-Level Characteristic Tone Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech-based emotion recognition systems face challenges in accurately recognizing emotions due to language, individual, and regional differences, leading to false recognition and a vicious cycle of user frustration, especially in interactive systems like automatic telephone answering systems and robots.
Innovation Solution
An emotion recognition apparatus that detects characteristic tones related to specific emotions in phonemes, using a characteristic tone detection unit, speech recognition unit, and emotion judgment unit, which computes a characteristic tone occurrence indicator to accurately recognize emotions without being affected by language or regional differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech-based emotion recognition methods are used, then the system can process user speech, but the recognition accuracy deteriorates due to language, individual, and regional differences
Solution Approach 1:
The patent segments speech into phoneme units and analyzes characteristic tones at the phoneme level rather than processing entire sentences. This segmentation allows the system to identify emotion-indicative patterns in individual phonemes while using language-specific phoneme databases to adapt to different languages and regions, thereby resolving the contradiction between recognition accuracy and language adaptability
Solution Approach 2:
The patent applies different processing strategies to different parts of the speech signal. Specifically, it identifies and weights characteristic tones in phonemes differently based on their emotional significance, using language-specific phoneme databases to handle regional variations. This local quality approach enables accurate emotion recognition across multiple languages by adapting the analysis to local linguistic characteristics
2Measurement precision
If the system requests re-input after false recognition, then it attempts to improve recognition accuracy, but user frustration increases creating a vicious cycle
Solution Approach 1:
The patent implements feedback by continuously monitoring characteristic tones during speech input and adjusting the recognition process in real-time. When emotion-indicative characteristic tones are detected, the system adapts its recognition criteria accordingly, reducing false rejections and the need for re-input requests. This feedback mechanism breaks the vicious cycle by improving accuracy before user frustration escalates
Solution Approach 2:
The patent performs preliminary analysis of characteristic tones in phonemes before final speech recognition determination. By detecting emotional states and adjusting recognition parameters in advance, the system prepares appropriate recognition criteria that account for the user's emotional state, thereby preventing false recognition and reducing the need for re-input requests that cause frustration
Data Source
AI summary
An emotion recognition apparatus performs accurate and stable speech-based emotion recognition, irrespective of individual, regional, and language differences of prosodic information. The emotion recognition apparatus includes: a speech recognition unit which recognizes types of phonemes included in the input speech; a characteristic tone detection unit which detects a characteristic tone that relates to a specific emotion, in the input speech; a characteristic tone occurrence indicator computation unit which computes a characteristic tone occurrence indicator for each of the phonemes, based on the types of the phonemes recognized by the speech recognition unit, the characteristic tone occurrence indicator relating to an occurrence frequency of the characteristic tone; and an emotion judgment unit which judges an emotion of the speaker in a phoneme at which the characteristic tone occurs in the input speech, based on the characteristic tone occurrence indicator computed by the characteristic tone occurrence indicator computing unit.


