Context-Specific Hot Words for Automated Assistant Invocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated assistants require users to utter specific 'hot words' to invoke them, which can be cumbersome, especially in scenarios like setting timers or controlling media playback.
Innovation Solution
Implementing 'dynamic' and 'context-specific' hot words that allow automated assistants to intelligently listen for relevant hot words based on the current context, expanding or altering their hot word vocabulary temporarily.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If the automated assistant operates in a limited hot word listening state, then power consumption and computing resources are conserved, but users must utter specific long tail hot words to invoke the assistant which can be cumbersome
Solution Approach 1:
The patent implements dynamic hot word listening by transitioning the automated assistant between different listening states (limited hot word listening state and continued listening state) based on contextual conditions. The system dynamically adjusts its listening behavior - remaining in limited state to conserve power, or transitioning to continued listening state when contextual conditions are met, allowing users to issue commands without re-invoking the assistant.
Solution Approach 2:
The system changes the operational parameters of the automated assistant by modifying the set of hot words being listened for. In the limited hot word listening state, only a first set of hot words is monitored. When contextual conditions are satisfied, the system transitions to listening for a second set of hot words that is different from the first set, thereby adapting the listening parameters to the current context without requiring full speech-to-text processing.
2Adaptability or versatility
If the automated assistant performs comprehensive speech-to-text processing for all utterances, then all user inputs can be interpreted, but power and computing resources are wasted
Solution Approach 1:
The patent applies partial action by implementing selective speech-to-text processing. Instead of processing all utterances comprehensively, the system performs STT processing only on a subset of utterances that meet specific contextual conditions. This partial processing approach maintains necessary response capability while avoiding energy waste on unnecessary processing of other utterances.
Solution Approach 2:
The patent segments the listening and processing functions into distinct modes. The automated assistant segments its operation into: (1) limited hot word listening state with minimal processing, and (2) continued listening state with selective STT processing. This segmentation allows the system to maintain versatility when needed while conserving energy during normal operation.
3Ease of operation
If the automated assistant transitions to continued listening mode, then users can issue subsequent commands without re-invoking, but the assistant may perform excessive STT processing wasting resources
Solution Approach 1:
The system uses feedback from contextual conditions to determine whether to transition to continued listening mode. Specific contextual conditions must be satisfied before the assistant transitions from limited hot word listening to continued listening state. This feedback mechanism ensures that continued listening is activated only when appropriate, preventing excessive resource consumption while maintaining ease of operation when contextual conditions are met.
Data Source
AI summary
Techniques are described herein for enabling the use of “dynamic” or “context-specific” hot words for an automated assistant. In various implementations, an automated assistant may be operated at least in part on a computing device. Audio data captured by a microphone may be monitored for default hot word(s). Detection of one or more of the default hot words may trigger transition of the automated assistant from a limited hot word listening state into a speech recognition state. Transition of the computing device into a given state may be detected, and in response, the audio data captured by the microphone may be monitored for context-specific hot word(s), in addition to or instead of the default hot word(s). Detection of the context-specific hot word(s) may trigger the automated assistant to perform a responsive action associated with the given state, without requiring detection of default hot word(s).


