Voice Assistant Listening Windows Adjusted by Breathing Patterns
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice assistants require users to repeatedly activate the microphone by saying wake-up words or pressing buttons, leading to increased latency and inability to predict user intentions during pauses, resulting in incomplete utterances.
Innovation Solution
A system that analyzes user breathing patterns and non-speech artifacts to dynamically adjust the listening time of voice assistants, allowing for continuous listening based on user intentions without requiring additional wake-up commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the voice assistant uses a static listening time, then the device structure is simple and energy consumption is low, but the user experience deteriorates due to incomplete utterance capture and increased latency
Solution Approach 1:
The patent applies dynamics by transitioning from a static listening time to a dynamic listening time that adapts based on user breathing patterns. The system continuously monitors breathing characteristics and adjusts the listening duration in real-time, allowing the voice assistant to extend or shorten its listening window according to the user's speech rhythm and pauses, thereby capturing complete utterances without requiring complex manual configuration
Solution Approach 2:
The system implements feedback by using the detected breathing pattern as input to continuously adjust the listening time. The breathing pattern detection provides real-time feedback about the user's speech state, and this feedback loop enables the system to optimize the listening duration dynamically, improving utterance capture while maintaining adaptive complexity only when needed
2Ease of operation
If the voice assistant requires wake-up words for each command, then energy consumption is reduced and false activations are minimized, but user convenience deteriorates due to repeated activation requirements
Solution Approach 1:
The system applies preliminary action by detecting breathing patterns during pauses to predict the user's intention to speak again before the wake-up word is needed. By analyzing the breathing characteristics that precede speech, the system proactively prepares to activate the microphone in advance, enabling seamless continuation of conversation without requiring the user to say wake-up words
Solution Approach 2:
The voice assistant performs self-service by automatically detecting when the user is about to speak through breathing pattern analysis and autonomously adjusting its listening state. The system serves itself by managing its own activation timing based on physiological cues, eliminating the need for manual wake-up commands while optimizing energy consumption through intelligent, context-aware activation
3Reliability
If the listening time is extended to capture complete utterances, then utterance capture completeness improves, but latency increases and response time deteriorates
Solution Approach 1:
The system resolves this contradiction by making the listening time dynamic rather than fixed. By continuously adapting the listening duration based on real-time breathing pattern analysis, the system extends listening only when the user shows signs of continuing speech (such as deep breathing before a long utterance) and maintains shorter listening windows when immediate response is needed, thereby optimizing both completeness and latency
Solution Approach 2:
The patent applies parameter changes by modifying the listening time parameter based on detected breathing characteristics. The system changes the listening duration parameter dynamically according to the user's physiological state, adjusting it to match the actual speech requirements without unnecessarily extending the listening window, thus reducing latency while ensuring complete capture
Data Source
AI summary
A method of adjusting a predefined listening time of a voice assistant device includes receiving an audio input; extracting at least one of a speech component and a non-speech artifact from the audio input; determining a user breathing pattern based on the at least one of the speech component and the non-speech artifact; identifying at least one attribute that impact the user breathing pattern based on at least one non-speech component, captured from an environment and the voice assistant device; determining, after detecting a pause in the audio input, whether a user's intention is to continue a conversation based on an analysis of the user breathing pattern and the at least one attribute; and dynamically adjusting the predefined listening time of the voice assistant device to continue listening for voice commands in the conversation based on a determination that the user's intention is to continue the conversation.


