Audio Analysis System for Emotion Detection via Non-Lexical Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice user interfaces face challenges in recognizing audio inputs due to poor quality, grammatical errors, and non-lexical aspects of speech, leading to mistriggers, and there is a need for systems to infer emotional states and intentions from non-lexical speech portions. Additionally, there is a lack of effective methods for identifying mental health states based on data collected from computing devices.

Innovation Solution

The development of methods and systems that analyze voice-based audio inputs by extracting non-lexical features such as articulation space, pitch, energy, and vocal effort, and using predictive models to identify emotional states and intentions, while also utilizing behavioral data from mobile devices to determine mental health states by processing sensor and usage data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems process only lexical content for voice commands, then processing speed is maintained, but accuracy deteriorates due to mistriggers caused by poor audio quality and non-lexical speech aspects

Engineering Contradiction:
Improveaudio input recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments audio input analysis into multiple independent components: lexical content processing, non-lexical feature extraction (pitch, energy, spectral features), and separate emotion/state detection modules. This allows the system to process different aspects of speech through dedicated pipelines, improving recognition accuracy without creating a monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements multi-functional processing where the same audio input is simultaneously analyzed for lexical meaning, emotional state, intent detection, and quality assessment. This universal approach allows a single system to handle multiple tasks that would otherwise require separate processing channels, improving accuracy while managing complexity through shared infrastructure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If voice user interfaces analyze non-lexical speech features to infer emotional states and intentions, then user understanding improves, but processing time increases

Engineering Contradiction:
Improveinformation completeness about user stateVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary extraction of non-lexical features (pitch contours, energy patterns, spectral characteristics) during the audio input phase itself, before full processing begins. These features are pre-computed and stored for subsequent emotion and intent analysis, reducing the computational burden during critical response phases and minimizing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediate representations of emotional state and user intent that serve as mediators between raw audio features and final system responses. These intermediate states compress complex non-lexical information into manageable categories that can be quickly processed and acted upon, reducing the time required for complete analysis while preserving essential information about user state.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If computing devices collect and analyze behavioral data from multiple sensors, then mental health state identification accuracy improves, but data privacy risks increase

Engineering Contradiction:
Improvemental health state detection accuracyVSAvoiddata privacy risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system extracts and processes only the specific behavioral features necessary for mental health state identification (such as speech patterns, usage behaviors, sensor data correlations) while leaving out unnecessary personal information. This selective extraction approach maintains detection accuracy by focusing on relevant indicators while reducing privacy risks by minimizing the collection and storage of sensitive personal data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent employs intermediary processing layers that transform raw behavioral data from multiple sensors into aggregated statistical representations before analysis. These intermediaries mask individual identifying characteristics while preserving the patterns necessary for mental health state detection, thereby maintaining accuracy while protecting user privacy through data anonymization and aggregation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11244698B2Systems and methods for identifying human emotions and/or mental health states based on analyses of audio inputs and/or behavioral data collected from computing devices
Publication Date: 2022.02.08 VERINT AMERICAS INC
  • US11244698B2 patent drawing
  • US11244698B2 patent drawing
  • US11244698B2 patent drawing

AI summary

Systems and methods are provided for analyzing voice-based audio inputs. A voice-based audio input associated with a user (e.g., wherein the voice-based audio input is a prompt or a command) is received and measures of one or more features are extracted. One or more parameters are calculated based on the measures of the one or more features. The occurrence of one or more mistriggers is identified by inputting the one or more parameters into a predictive model. Further, systems and methods are provided for identifying human mental health states using mobile device data. Mobile device data (including sensor data) associated with a mobile device corresponding to a user is received. Measurements are derived from the mobile device data and input into a predictive model. The predictive model is executed and outputs probability values of one or more symptoms associated with the user.