Adaptive Speech Activity Detection Threshold for Barge-In
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech dialogue systems face challenges in accurately and reliably detecting barge-in during speech prompts, leading to potential misclassification of speech prompts as speech input.
Innovation Solution
A method and apparatus that adjust a time-varying sensitivity threshold and utilize speaker information to differentiate between speech prompts and speech activity, ensuring the speech activity detector is active during speech prompts and adapting the sensitivity based on speaker behavior to minimize errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If the speech activity detector is activated during speech prompt output to enable barge-in detection, then the system can recognize user speech earlier, but the speech prompt itself may be erroneously classified as speech input
Solution Approach 1:
The patent applies dynamics by making the detection threshold time-varying rather than static. The threshold adapts dynamically based on whether a speech prompt is currently being output, allowing the system to be sensitive to user speech during prompts while remaining insensitive to the prompt itself. This resolves the contradiction by enabling continuous monitoring (reducing wait time) while maintaining detection accuracy through adaptive thresholding.
Solution Approach 2:
The patent changes the parameter of detection sensitivity by adjusting the threshold level based on the state of speech prompt output. When a prompt is detected, the threshold is raised to prevent false detection of the prompt as user speech. When no prompt is active, the threshold returns to normal sensitivity levels. This parameter adaptation enables reliable barge-in detection without false positives.
2Reliability
If a fixed high threshold is used to avoid misclassifying speech prompts as speech input, then false positives are reduced, but the system cannot detect barge-in during prompt output
Solution Approach 1:
Instead of using a fixed high threshold that prevents barge-in detection, the patent employs a dynamic threshold that adapts to the operational context. During speech prompt output, the threshold is elevated to avoid false positives. During periods when no prompt is active, the threshold decreases to normal levels, enabling timely detection of user speech and maintaining dialogue efficiency.
Solution Approach 2:
The system performs preliminary detection of speech prompt output and uses this information to pre-adjust the detection threshold before user speech may occur. By anticipating the need for different sensitivity levels based on prompt state, the system prepares the appropriate threshold configuration in advance, enabling both accurate differentiation and timely response.
3Reliability
If the speech activity detector remains inactive during speech prompt output to avoid false detection, then detection accuracy is maintained, but the user must wait for the prompt to finish before speaking
Solution Approach 1:
The patent resolves this contradiction by dynamically controlling the speech activity detector based on prompt state. The detector remains active throughout, but its sensitivity is modulated via threshold adjustment. During prompt output, the elevated threshold prevents false detection while the detector continues monitoring. When prompts end, the threshold decreases and the detector becomes fully sensitive, eliminating the need for users to wait for prompt completion.
Solution Approach 2:
The system uses feedback from speech prompt detection to continuously adjust the speech activity detection threshold. The prompt detection output feeds into the threshold control mechanism, creating a closed-loop system that automatically adapts sensitivity based on operational context. This feedback mechanism enables continuous monitoring without false positives, allowing immediate user response regardless of prompt state.
Data Source
AI summary
A method for detecting barge-in in a speech dialog system comprising determining whether a speech prompt is output by the speech dialog system, and detecting whether speech activity is present in an input signal based on a time-varying sensitivity threshold of a speech activity detector and/or based on speaker information, where the sensitivity threshold is increased if output of a speech prompt is determined and decreased if no output of a speech prompt is determined. If speech activity is detected in the input signal, the speech prompt may be interrupted or faded out. A speech dialog system configured to detect barge-in is also disclosed.


