Speech Emotion Analysis via Acoustic Space Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition and analysis techniques fail to effectively incorporate emotion as an integral component of human speech, limiting their ability to fully understand and interpret speech patterns.
Innovation Solution
A method and apparatus for analyzing speech that involves receiving an utterance, converting it into a speech signal, and comparing segments based on time and frequency to discriminate emotions using acoustic characteristics, creating an acoustic space to determine emotion states by measuring and comparing baseline characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If current voice recognition techniques are used, then language parsing and identification is achieved, but emotion analysis capability is lost
Solution Approach 1:
The speech signal is divided into segments based on time and frequency, allowing separate analysis of segmental and suprasegmental properties. This segmentation enables the system to capture both linguistic content and emotional expressions that would otherwise be lost in conventional voice recognition.
Solution Approach 2:
The patent introduces a new dimension of analysis by creating an acoustic space that incorporates both segmental and suprasegmental properties. This dimensional expansion allows the system to represent speech not only linguistically but also emotionally, resolving the contradiction between language parsing and emotion analysis.
2Measurement precision
If segmental and suprasegmental properties are analyzed, then emotion discrimination accuracy is improved, but system complexity increases
Solution Approach 1:
By dividing the speech signal into manageable segments based on time and frequency, the system can systematically analyze both segmental and suprasegmental properties. This segmentation approach organizes the complex analysis task into discrete, processable units, improving emotion discrimination while managing computational complexity.
Solution Approach 2:
The baseline is determined from acoustic characteristics of emotion categories before analyzing individual speech segments. This preliminary preparation of reference data enables more efficient and accurate emotion discrimination during actual speech analysis, reducing real-time processing complexity.
Data Source
AI summary
A method and apparatus for analyzing speech are provided. A method and apparatus for determining an emotion state of a speaker are provided, including providing an acoustic space having one or more dimensions, where each dimension corresponds to at least one baseline acoustic characteristic; receiving an utterance of speech by the speaker; measuring one or more acoustic characteristics of the utterance; comparing each of the measured acoustic characteristics to a corresponding baseline acoustic characteristic; and determining an emotion state of the speaker based on the comparison. An embodiment involves determining the emotion state of the speaker within one day of receiving the subject utterance of speech. An embodiment involves determining the emotion state of the speaker, where the emotion state of the speaker includes at least one magnitude along a corresponding at least one of the one or more dimensions within the acoustic space.


