Speech Emotion Identification via Differential Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech signal processing systems for emotion identification lack consistency and accuracy in recognizing emotions conveyed through speech signals, as different mechanisms exhibit varying levels of effectiveness.

Innovation Solution

A processor-implemented method and system that samples speech signals at a defined rate, splits them into frames, extracts differential features, and compares these features with an emotion recognition model to identify and associate emotions, using a specific technique of selecting one sample every M samples and calculating differences between adjacent samples to generate output frames for emotion identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional emotion identification mechanisms are used, then the system can identify emotions in speech signals, but the accuracy and consistency of emotion recognition varies and lacks reliability

Engineering Contradiction:
Improveemotion identification accuracyVSAvoidemotion recognition consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The speech signal is segmented into multiple frames of 20 milliseconds each, allowing differential features to be extracted from each frame independently. This segmentation enables consistent processing of speech signals while capturing temporal variations in emotional expression, thereby improving both accuracy and reliability of emotion identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies periodic sampling at a defined rate and processes each frame through consistent differential feature extraction. By using fixed frame lengths and regular sampling intervals, the system ensures reproducible and consistent emotion recognition results across different speech signals, enhancing reliability while maintaining accuracy.

Inventive Principle:
Principle #19Periodic action

2Measurement precision

If differential features are extracted from speech signals using the specified method, then emotion identification accuracy is enhanced, but the processing complexity increases

Engineering Contradiction:
Improveemotion identification accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Instead of processing every single sample, the method selects one sample in every M samples (where M=2) and calculates differences only between adjacent selected samples. This partial processing approach reduces computational complexity while still capturing sufficient differential features to maintain high emotion identification accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system extracts only the essential differential features from speech signals by calculating differences between adjacent samples and selecting specific samples at regular intervals. This extraction of critical features rather than processing the entire signal reduces complexity while preserving the information necessary for accurate emotion identification.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11227624B2Method and system using successive differences of speech signals for emotion identification
Publication Date: 2022.01.18 TATA CONSULTANCY SERVICES LTD
  • US11227624B2 patent drawing
  • US11227624B2 patent drawing
  • US11227624B2 patent drawing

AI summary

This disclosure relates generally to speech signal processing, and more particularly to a method and system for processing speech signal for emotion identification. The system processes a speech signal collected as input, during which a plurality of differential features corresponding to a plurality of frames of the speech signal are extracted. Further, the differential features are compared with an emotion recognition model to identify at least one emotion matching the speech signal, and then the at least one emotion is associated with the speech signal.