Speech Analysis for OSA Severity Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for diagnosing obstructive sleep apnea (OSA) are invasive, labor-intensive, and costly, particularly polysomnography (PSG), which requires patients to be asleep, and there is a need for simpler, cost-effective approaches to estimate OSA severity in wakeful patients, especially considering dialectal and linguistic variations.
Innovation Solution
A system and method that combine acoustic short-term features, long-term features, and features of sustained vowels from speech signals to estimate the apnea-hypopnea index (AHI) using statistical learning and speech analysis, processing audio recordings to generate a fused AHI estimate through pre-defined computing models and machine learning techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If polysomnography (PSG) is used to diagnose OSA, then measurement precision is improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent extracts the essential diagnostic function from the complex PSG system by isolating speech signal analysis as a standalone screening method. Speech descriptors are extracted from normal conversation recordings without requiring the full PSG apparatus, thereby simplifying the diagnostic process while maintaining OSA detection capability
Solution Approach 2:
The patent replaces the mechanical and physiological measurement system of PSG with an acoustic signal processing system. Instead of measuring respiratory flows, oxygen saturation, and brain waves through complex sensors, the system substitutes speech acoustic features (formant frequencies, spectral characteristics) as proxies for upper airway anatomy and function
2Measurement precision
If polysomnography (PSG) is used to diagnose OSA, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent performs preliminary speech analysis during wakeful hours when patients are naturally speaking, capturing speech samples before the actual sleep study. This preliminary action allows the system to pre-process and analyze speech descriptors, reducing the time required during the actual PSG or enabling screening without requiring the patient to complete a full overnight study
Solution Approach 2:
The patent creates a simplified copy of the diagnostic process by using speech recordings (which can be obtained during routine clinic visits) as a substitute for the full PSG procedure. The speech-based AHI estimation replicates the essential diagnostic function without requiring the time-consuming overnight hospital stay
3Ease of operation
If speech analysis is used to estimate AHI, then ease of operation is improved, but measurement precision deteriorates
Solution Approach 1:
The patent merges multiple speech descriptors (formant frequencies F1, F2, F3; spectral characteristics; temporal features) from different aspects of speech production into a unified AHI estimation model. By combining these diverse acoustic features, the system compensates for the limitations of any single descriptor and improves overall measurement precision
Solution Approach 2:
The patent transforms speech signals from the time domain to the frequency domain through spectral analysis, extracting formant frequencies and spectral characteristics that are not apparent in the raw waveform. This parameter transformation reveals hidden diagnostic information in speech that correlates with upper airway anatomy and OSA severity
4Loss of substance
If speech analysis is used to estimate AHI, then loss of substance is reduced, but reliability deteriorates
Solution Approach 1:
The patent develops a universal speech analysis method that functions across different languages, dialects, and speaking styles by focusing on fundamental acoustic properties of human speech (formants, spectral characteristics) that are language-independent. This universality allows the system to maintain reliability across diverse populations without requiring language-specific calibration
Data Source
AI summary
Provided herein is a method and system for the estimation of apnea-hypopnea index (AHI), as an indicator for Obstructive sleep apnea (OSA) severity, by combining speech descriptors from three separate and distinct speech signal domains. These domains include the acoustic short-term features (STF) of continuous speech, the long-term features (LTF) of continuous speech, and features of sustained vowels (SVF). Combining these speech descriptors may provide the ability to estimate the severity of OSA using statistical learning and speech analysis approaches.


