Acoustic Feature Extraction for Mood State Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for monitoring mood changes in individuals with bipolar disorder are resource-intensive, costly, and lack reliable, objective biomarkers, as they rely on clinical observations and self-reported data, which are incomplete and variable, limiting the development of accurate mood recognition systems.
Innovation Solution
A method and system that use a mobile device to record audio samples, extract acoustic features, and generate emotion values using a trained machine learning model to predict mood states, providing a cost-efficient and objective means of monitoring mood changes by analyzing speech patterns and emotional expressions in natural conversations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If clinical observations and self-reported data are used for mood monitoring, then mood changes can be tracked, but the methods are resource-intensive, costly, and lack reliability due to variable clinical observations and incomplete self-reports
Solution Approach 1:
The patent replaces manual clinical observations and self-reported data collection with an automated acoustic analysis system. Machine learning models analyze speech audio signals to objectively detect mood states, substituting the mechanical processes of clinical evaluation with automated computational analysis. This reduces human resource requirements while improving reliability through consistent, objective measurements.
Solution Approach 2:
The system enables self-monitoring of mood states through automated analysis of speech patterns. The machine learning model processes audio data from mobile devices to generate mood assessments without requiring clinical intervention, allowing the system to serve itself and the patient continuously without external resource input.
2Measurement precision
If direct speech-to-mood mapping is attempted, then mood recognition can be achieved, but the complexity of the speech signal makes accurate recognition challenging
Solution Approach 1:
The patent introduces acoustic features as an intermediary layer between raw speech signals and mood state classification. The system extracts specific acoustic characteristics (pitch, intensity, spectral features) that serve as mediators, translating complex speech variability into standardized features that the machine learning model can process effectively. This intermediary representation simplifies the mapping process while maintaining accuracy.
Solution Approach 2:
The system transforms the complex speech signal into a reduced set of meaningful acoustic parameters. By changing the representation from raw audio waves to extracted acoustic features (frequency, amplitude, temporal characteristics), the system simplifies the input data while preserving the essential information needed for mood recognition, making the processing more manageable and accurate.
3Loss of information
If self-reported diagnosis data from mobile devices or social media are used, then mood information can be collected, but the data are often incomplete or misleading
Solution Approach 1:
The system automatically collects mood-related information through speech analysis without requiring active patient participation or self-reporting. The mobile device's microphone continuously captures speech, and the machine learning model automatically processes this data to generate mood assessments, eliminating the need for patients to manually input data while ensuring complete and accurate information collection.
Solution Approach 2:
The patent replaces the manual self-reporting mechanism with automated acoustic sensing. Instead of relying on patients to consciously report their mood states (which can be incomplete or misleading), the system uses the microphone to capture speech patterns and the machine learning model to objectively infer mood, substituting the subjective reporting process with an automated sensory-mechanical system.
Data Source
AI summary
A method of predicting a mood state of a user may include recording an audio sample via a microphone of a mobile computing device of the user based on the occurrence of an event, extracting a set of acoustic features from the audio sample, generating one or more emotion values by analyzing the set of acoustic features using a trained machine learning model, and determining the mood state of the user, based on the one or more emotion values. In some embodiments, the audio sample may be ambient audio recorded periodically, and/or call data of the user recorded during clinical calls or personal calls.


