Glottal Waveform Feature Extraction for Depression Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for diagnosing depression are time-consuming, qualitative, and often fail to detect depression in its early stages due to limited availability of mental health professionals, leading to untreated cases and potential suicides.

Innovation Solution

A method and device that analyze natural speech signals to classify mental states by extracting glottal waveform features and comparing them with pre-determined parameters using a classifier, enabling real-time diagnosis of depression and suicidal risk.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated speech analysis is implemented, then diagnostic speed and accessibility are improved, but measurement precision and reliability may be compromised compared to expert clinical judgment

Engineering Contradiction:
Improvediagnostic speedVSAvoiddiagnostic accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system creates a computational model that replicates expert clinical judgment by training on speech samples from patients with known diagnoses. The classifier learns to copy the diagnostic patterns that expert practitioners use, enabling automated systems to achieve diagnostic accuracy comparable to human experts while providing rapid assessment at scale.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary screening and assessment through automated speech analysis before referring cases to expert clinicians. By pre-processing and filtering cases using the trained classifier, the system prepares and prioritizes cases in advance, enabling experts to focus their attention on cases that require their specialized judgment while rapidly assessing many other cases automatically.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If comprehensive speech feature analysis is performed, then diagnostic accuracy is improved, but device complexity and computational requirements increase

Engineering Contradiction:
Improvediagnostic accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts and analyzes specific glottal waveform features from speech signals that have been identified as most relevant for depression detection. Rather than analyzing all possible speech features, the system focuses on extracting glottal waveform characteristics through pitch contour analysis and spectral features, simplifying the overall system while maintaining diagnostic accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms raw speech signals into specific parameter representations including fundamental frequency contours, spectral features, and glottal waveform characteristics. By changing the parameter space from raw audio to extracted features, the system reduces computational complexity while preserving the diagnostic information needed for accurate classification.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9058816B2Emotional and/or psychiatric state detection
Publication Date: 2015.06.16 RMIT UNIVERSITY
  • US9058816B2 patent drawing
  • US9058816B2 patent drawing
  • US9058816B2 patent drawing

AI summary

Mental state of a person is classified in an automated manner by analysing natural speech of the person. A glottal waveform is extracted from a natural speech signal. Pre-determined parameters defining at least one diagnostic class of a class model are retrieved, the parameters determined from selected training glottal waveform features. The selected glottal waveform features are extracted from the signal. Current mental state of the person is classified by comparing extracted glottal waveform features with the parameters and class model. Feature extraction from a glottal waveform or other natural speech signal may involve determining spectral amplitudes of the signal, setting spectral amplitudes below a pre-defined threshold to zero and, for each of a plurality of sub bands, determining an area under the thresholded spectral amplitudes, and deriving signal feature parameters from the determined areas in accordance with a diagnostic class model.