Voice Activity Detection Customization for Dysarthric Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dialog systems face challenges in accurately detecting voice activity in dysarthric speech, which is characterized by poor articulation, disfluencies, and atypical acoustic properties, affecting the performance of voice activity detection (VAD) in multimodal dialog applications for ALS patients and others with dysarthria.

Innovation Solution

A cloud-based dialog agent that customizes voice activity detection by identifying spans of speech and non-speech in an audio stream, adjusting configurable parameters such as minimum signal strength, background noise adjustment, and speech and non-speech thresholds to optimize user experience for dysarthric speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If standard voice activity detection parameters are used for non-disordered speech, then the system performs well for healthy speakers, but the detection accuracy substantially decreases for dysarthric speakers

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidadaptability to different speech pathologies
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by systematically adjusting VAD configuration parameters (minimum signal strength, background noise adjustment factors, speech/non-speech thresholds, start speech time threshold, and end speech time threshold) to optimize detection accuracy for dysarthric speech characteristics, thereby resolving the contradiction between maintaining high detection accuracy and adapting to different speech pathologies

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If voice activity detection is optimized for dysarthric speech characteristics, then detection accuracy for dysarthric speakers improves, but false positives and false negatives increase for non-dysarthric speech

Engineering Contradiction:
Improvedetection accuracy for dysarthric speakersVSAvoidfalse positive and false negative rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies dynamics by implementing adaptive parameter adjustment that dynamically modifies VAD configuration based on detected speech characteristics. The system monitors speech signals and adjusts parameters in real-time to match the speaker's pathology type, thereby maintaining high detection accuracy while minimizing false positives and false negatives across different speaker populations

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If multiple configurable parameters are adjusted to accommodate different speech pathologies, then user experience across diverse cohorts improves, but system complexity increases

Engineering Contradiction:
Improvecustomization for different cohorts and tasksVSAvoidnumber of configurable VAD parameters
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-configuring optimized parameter sets for different speech pathology types and task categories before actual use. The system includes pre-established parameter configurations that can be automatically selected based on the detected speaker cohort and task type, thereby achieving high adaptability without requiring complex real-time parameter tuning during operation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12300227B2Customizing computer generated dialog for different pathologies
Publication Date: 2025.05.13 MODALITY AI
  • US12300227B2 patent drawing
  • US12300227B2 patent drawing
  • US12300227B2 patent drawing

AI summary

A computer-generated dialog session is customized for a user having a pathology characterized at least in part by a speech pathology. The user's speech is analyzed for spans of speech in which the starts and ends of the spans satisfy predetermined thresholds of time. Customization occurs by altering at least one of the following configurable parameters: (a) a threshold minimum signal strength of speech (dB) to consider as the start of the span of speech; (b) an adjustment factor by which signal strengths of background noise increases between consecutive spans of speech; (c) a threshold between signal strength during the span of speech and signal strength during the span of non-speech; (d) a start speech time threshold; and (e) an end speech time threshold.