Speech Emotion Analysis via Acoustic Space Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice recognition and analysis techniques fail to effectively incorporate emotion as an integral component of human speech, limiting their ability to fully understand and interpret speech patterns.

Innovation Solution

A method and apparatus for analyzing speech that involves receiving an utterance, converting it into a speech signal, and comparing segments based on time and frequency to discriminate emotions using acoustic characteristics, creating an acoustic space to determine emotion states by measuring and comparing baseline characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If current voice recognition techniques are used, then language parsing and identification is achieved, but emotion analysis capability is lost

Engineering Contradiction:
Improveemotion informationVSAvoidspeech analysis capability
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The speech signal is divided into segments based on time and frequency, allowing separate analysis of segmental and suprasegmental properties. This segmentation enables the system to capture both linguistic content and emotional expressions that would otherwise be lost in conventional voice recognition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of analysis by creating an acoustic space that incorporates both segmental and suprasegmental properties. This dimensional expansion allows the system to represent speech not only linguistically but also emotionally, resolving the contradiction between language parsing and emotion analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If segmental and suprasegmental properties are analyzed, then emotion discrimination accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveemotion discrimination accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

By dividing the speech signal into manageable segments based on time and frequency, the system can systematically analyze both segmental and suprasegmental properties. This segmentation approach organizes the complex analysis task into discrete, processable units, improving emotion discrimination while managing computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The baseline is determined from acoustic characteristics of emotion categories before analyzing individual speech segments. This preliminary preparation of reference data enables more efficient and accurate emotion discrimination during actual speech analysis, reducing real-time processing complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8788270B2Apparatus and method for determining an emotion state of a speaker
Publication Date: 2014.07.22 UNIV OF FLORIDA RESEARCH FOUNDATION INC
  • US8788270B2 patent drawing
  • US8788270B2 patent drawing
  • US8788270B2 patent drawing

AI summary

A method and apparatus for analyzing speech are provided. A method and apparatus for determining an emotion state of a speaker are provided, including providing an acoustic space having one or more dimensions, where each dimension corresponds to at least one baseline acoustic characteristic; receiving an utterance of speech by the speaker; measuring one or more acoustic characteristics of the utterance; comparing each of the measured acoustic characteristics to a corresponding baseline acoustic characteristic; and determining an emotion state of the speaker based on the comparison. An embodiment involves determining the emotion state of the speaker within one day of receiving the subject utterance of speech. An embodiment involves determining the emotion state of the speaker, where the emotion state of the speaker includes at least one magnitude along a corresponding at least one of the one or more dimensions within the acoustic space.