Voice-Based System for Objective Emotional Distress Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mental health treatment methods lack an objective way to quantify the effectiveness of therapy sessions for patients, particularly children and adolescents, due to their difficulty in articulating emotions, leading to unclear assessment of emotional distress and treatment outcomes.
Innovation Solution
A voice-based system that uses machine learning models to analyze audio recordings and predict emotional distress severity by identifying archetypal emotions such as anger, fear, happiness, and sadness, leveraging prosodic and acoustic features, and integrating with clinical diagnoses to provide an objective measure of emotional trauma.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If therapists rely on patient self-reporting for emotional assessment, then the assessment process is simple and non-intrusive, but the measurement precision is poor due to patients' difficulty in articulating emotions
Solution Approach 1:
The patent replaces the mechanical/verbal communication system with an acoustic analysis system. Instead of relying on patients to verbally articulate their emotions, the system uses audio recording and acoustic feature extraction to objectively measure emotional distress, substituting the human communication mechanism with an automated acoustic measurement system.
Solution Approach 2:
The patent introduces an intermediary system between the patient and the assessment process. The acoustic analysis system acts as a mediator that captures emotional information through audio recordings and processes it through machine learning models, eliminating the need for direct verbal self-reporting while maintaining non-intrusive assessment.
2Productivity
If therapists conduct detailed conversations to assess emotional distress, then the assessment can capture nuanced emotional states, but the measurement process becomes time-consuming and labor-intensive
Solution Approach 1:
The system performs preliminary acoustic feature extraction and emotional classification during or after therapy sessions, rather than requiring therapists to conduct lengthy assessments during sessions. The audio recording and analysis can be performed in advance or immediately after, freeing up therapist time while maintaining assessment quality.
Solution Approach 2:
The acoustic analysis system performs self-service by automatically processing audio recordings and generating emotional distress assessments without requiring therapist intervention for the measurement process itself. The machine learning model independently analyzes acoustic features and provides diagnostic insights, reducing the time and labor required from therapists.
3Reliability
If the system uses complex machine learning models to analyze emotional patterns, then the prediction accuracy improves, but the device complexity and computational requirements increase
Solution Approach 1:
The patent segments the complex emotional analysis process into distinct modular components: audio recording, acoustic feature extraction, machine learning classification, and diagnostic output. This segmentation allows each component to be optimized independently and simplifies the overall system architecture while maintaining high prediction reliability through specialized processing at each stage.
Data Source
AI summary
Systems and methods of related to a voice-based system used to determine the severity of emotional distress within an audio recording of an individual is provided. In one non-limiting example, a system comprises a computing device that is configured to receive an audio sample that includes an utterance of a user. Feature extraction is performed on the audio sample to extract a plurality of acoustic emotion features using a base model. Emotion level predictions are generated for an emotion type based at least in part on the acoustic emotion features provided to an emotion specific model. An emotion classification for the audio sample is determined based on the emotion level predictions. The emotion classification comprises the emotion type and a level associated with the emotion type.


