Character-Level Emotion Detection Network for Speech Sentiment Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals are often unaware of the emotional state they convey through their speech, which can impact their health and interactions with others, and there is a lack of feasible methods to continuously provide feedback on this emotional state.

Innovation Solution

A system and method that processes speech to determine sentiment data indicative of emotional states by converting audio data into character data using a character-level emotion detection network, providing output to the user through a user interface, allowing for self-assessment and behavioral modification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech is continuously monitored to provide emotional state feedback, then users can better understand and manage their emotional state, but this requires complex processing systems and raises privacy concerns

Engineering Contradiction:
Improveemotional state detection accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential acoustic features (pitch, energy, spectral characteristics) needed for emotion detection rather than processing the entire speech signal or audio stream. This extraction approach maintains detection accuracy while significantly reducing computational complexity and data processing requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces an intermediary processing layer that converts raw audio signals into emotional state indicators through acoustic feature analysis. This intermediary layer acts as a mediator between the complex raw speech data and the simple emotional state output, making the system more manageable and interpretable.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If detailed acoustic features are analyzed to improve emotion detection accuracy, then sentiment analysis becomes more precise, but processing time and computational resources increase

Engineering Contradiction:
Improvesentiment analysis accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The speech signal is segmented into short frames (e.g., 20-50 milliseconds) and analyzed independently for acoustic features. This segmentation allows the system to process speech in manageable chunks, reducing the computational burden on each frame while maintaining overall detection accuracy through continuous analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system analyzes only the most discriminative acoustic features (pitch contours, energy variations, spectral moments) rather than computing all possible speech parameters. This partial analysis approach focuses computational resources on the features that most strongly correlate with emotional states, achieving good accuracy with reduced processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11869535B1Character-level emotion detection
Publication Date: 2024.01.09 AMAZON TECH INC
  • US11869535B1 patent drawing
  • US11869535B1 patent drawing
  • US11869535B1 patent drawing

AI summary

Described is a system and method that determines character sequences from speech, without determining the words of the speech, and processes the character sequences to determine sentiment data indicative of emotional state of a user that output the speech. The emotional state may then be presented or provided as an output to the user.