Voice Activity Detection Customization for Dysarthric Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing dialog systems face challenges in accurately detecting voice activity in dysarthric speech, which is characterized by poor articulation, disfluencies, and atypical acoustic properties, affecting the performance of voice activity detection (VAD) in multimodal dialog applications for ALS patients and others with dysarthria.
Innovation Solution
A cloud-based dialog agent that customizes voice activity detection by identifying spans of speech and non-speech in an audio stream, adjusting configurable parameters such as minimum signal strength, background noise adjustment, and speech and non-speech thresholds to optimize user experience for dysarthric speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If standard voice activity detection parameters are used for non-disordered speech, then the system performs well for healthy speakers, but the detection accuracy substantially decreases for dysarthric speakers
Solution Approach 1:
The patent applies parameter changes by systematically adjusting VAD configuration parameters (minimum signal strength, background noise adjustment factors, speech/non-speech thresholds, start speech time threshold, and end speech time threshold) to optimize detection accuracy for dysarthric speech characteristics, thereby resolving the contradiction between maintaining high detection accuracy and adapting to different speech pathologies
2Measurement precision
If voice activity detection is optimized for dysarthric speech characteristics, then detection accuracy for dysarthric speakers improves, but false positives and false negatives increase for non-dysarthric speech
Solution Approach 1:
The patent applies dynamics by implementing adaptive parameter adjustment that dynamically modifies VAD configuration based on detected speech characteristics. The system monitors speech signals and adjusts parameters in real-time to match the speaker's pathology type, thereby maintaining high detection accuracy while minimizing false positives and false negatives across different speaker populations
3Adaptability or versatility
If multiple configurable parameters are adjusted to accommodate different speech pathologies, then user experience across diverse cohorts improves, but system complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-configuring optimized parameter sets for different speech pathology types and task categories before actual use. The system includes pre-established parameter configurations that can be automatically selected based on the detected speaker cohort and task type, thereby achieving high adaptability without requiring complex real-time parameter tuning during operation
Data Source
AI summary
A computer-generated dialog session is customized for a user having a pathology characterized at least in part by a speech pathology. The user's speech is analyzed for spans of speech in which the starts and ends of the spans satisfy predetermined thresholds of time. Customization occurs by altering at least one of the following configurable parameters: (a) a threshold minimum signal strength of speech (dB) to consider as the start of the span of speech; (b) an adjustment factor by which signal strengths of background noise increases between consecutive spans of speech; (c) a threshold between signal strength during the span of speech and signal strength during the span of non-speech; (d) a start speech time threshold; and (e) an end speech time threshold.


