Voice Command Stop Detection Without Speech Pause
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice command systems require users to pause or cease speaking to indicate the end of a voice command, which is inconvenient, especially in social settings like vehicles, where it disrupts conversation and requires unnecessary pauses.
Innovation Solution
A computing device detects a spoken utterance stop event that is not a pause or cessation, allowing users to continue speaking without interruption, by recognizing changes in direction, speech pattern, context, or specific phrases to determine when a voice command is complete.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system requires users to pause or cease speaking to indicate the end of a voice command, then the system can accurately detect the end of the command, but the user experience is disrupted and conversation flow is interrupted
Solution Approach 1:
The patent replaces the mechanical pause-based stop mechanism with an acoustic analysis system that detects speech characteristics. The system uses speech recognition to analyze phoneme sequences, silence durations, and speech patterns to automatically determine when a command ends, eliminating the need for users to manually pause or cease speaking.
Solution Approach 2:
The patent introduces speech recognition technology as an intermediary between the user's speech and the command execution system. This intermediary analyzes the speech stream in real-time, detecting subtle acoustic cues and patterns to identify command boundaries without requiring explicit user actions like pausing or stopping speech.
2Ease of operation
If the system continuously monitors speech to detect command end events, then the user can speak naturally without pauses, but the system complexity increases
Solution Approach 1:
The patent applies partial monitoring by focusing speech analysis only on specific detect events rather than continuous full speech processing. The system monitors for particular speech patterns, silence durations, and phoneme sequences that indicate command boundaries, rather than analyzing every aspect of the speech stream continuously.
Solution Approach 2:
The patent segments the speech stream into detect events with specific characteristics. Each detect event is analyzed independently for speech patterns, silence durations, and phoneme sequences. This segmentation allows the system to process speech in manageable units rather than as a continuous complex stream.
3Ease of manufacture
If the system uses traditional pause detection to identify command end, then the implementation is simple, but the system cannot handle continuous speech in social settings
Solution Approach 1:
The patent changes the detection parameters from simple pause duration to multiple speech characteristics including phoneme sequences, silence durations, speech patterns, and context analysis. This allows the system to adapt to various social settings and speech styles while maintaining accurate command boundary detection.
Solution Approach 2:
The patent implements dynamic detection that adapts to different speech contexts. The system adjusts its detection criteria based on the speech stream characteristics, allowing it to handle both formal commands and casual conversation in social settings. The detection parameters are not fixed but dynamically adjusted based on the observed speech patterns.
Data Source
AI summary
Speech recognition of a stream of spoken utterances is initiated. Thereafter, a spoken utterance stop event to stop the speech recognition is detected, such as in in relation to the stream. The spoken utterance stop event is other than a pause or cessation in the stream of spoken utterances. In response to the spoken utterance stop event being detected, the speech recognition of the stream of spoken utterances is stopped, while the stream of spoken utterances continues. After stopping the speech recognition of the stream of spoken utterances has been stopped, an action is caused to be performed that corresponds to the spoken utterances from a beginning of the stream through and until the spoken utterance stop event.


