Dynamic Listening Timeout for ASR Stutter Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition (ASR) systems use statically set listening timeouts, which do not accommodate variations in user speech patterns, potentially cutting off users who require more time to speak, especially those with stuttered or filler speech.
Innovation Solution
A dynamically-adjustable listening timeout system that processes received speech to determine the presence of insignificant utterances, such as stuttered or filler speech, and adjusts the timeout duration based on this analysis, allowing more time for users to complete their speech inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a static listening timeout is used, then the system operation is simple, but the system cannot accommodate variations in user speech patterns
Solution Approach 1:
The patent implements a dynamic listening timeout that automatically adjusts its duration based on real-time analysis of the user's speech characteristics. The system transitions from a fixed timeout value to a variable timeout that adapts to each user's speech patterns, particularly accommodating users with stuttered or filler speech without requiring manual configuration or complex user profiles.
Solution Approach 2:
The system performs self-adjustment by automatically analyzing the incoming speech stream to detect insignificant utterances and dynamically modifying the timeout duration accordingly. This eliminates the need for external configuration, user profiling, or manual intervention, allowing the system to serve itself in adapting to different users' speech patterns.
2Reliability
If a static listening timeout is used, then the system is easy to operate, but users requiring more time to speak are cut off
Solution Approach 1:
The system continuously monitors the speech stream for indicators of insignificant utterances (such as stuttering patterns or filler words) and uses this feedback to dynamically adjust the timeout duration. This closed-loop feedback mechanism ensures that the timeout is reliably adapted to the user's actual speech needs while maintaining automatic operation without user intervention.
Solution Approach 2:
The system performs preliminary analysis of the speech stream during the listening period to detect speech characteristics before the timeout expires. By identifying signs of stuttered or filler speech early in the utterance, the system can proactively extend the timeout to ensure complete capture of the user's intended speech, preventing premature cutoff.
3Measurement precision
If the listening timeout is extended for users with stuttered speech, then speech recognition accuracy improves, but the system cannot distinguish between different speech types
Solution Approach 1:
The system applies different timeout adjustment strategies based on the local characteristics of the detected speech. Instead of using a uniform approach for all speech, the system identifies specific patterns (stuttered speech, filler speech, incomprehensible speech) and applies targeted timeout extensions appropriate to each speech type, optimizing recognition accuracy for each case.
Solution Approach 2:
The system dynamically changes the timeout parameter based on the detected speech characteristics. By monitoring speech parameters such as pause duration, repetition frequency, and utterance structure, the system adjusts the timeout duration to match the specific requirements of different speech types, transitioning from a single fixed parameter to a dynamically modified parameter set.
Data Source
AI summary
A system and method of automated speech recognition using a dynamically-adjustable listening timeout. The method includes: receiving speech signals representing a first speech segment during a first speech listening period; during the first speech listening period, processing the received speech signals representing the first speech segment to determine whether the first speech segment includes one or more insignificant utterances; adjusting a listening timeout in response to the determination of whether the first speech segment includes one or more insignificant utterances; listening for subsequent received speech using the adjusted listening timeout; and performing automatic speech recognition on the received speech signals and/or the subsequently received speech signals.


