Automated Assistant Commands Without Wake Words for Display Content
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated assistants require invocation phrases or words, which prolong interactions and inefficiently consume resources, especially when rendering content like audio or text-to-speech operations, and do not allow invocation-free utterances to perform operations.
Innovation Solution
An automated assistant that responds to a dynamically updated set of words or phrases based on display content, allowing invocation-free utterances to perform operations without requiring a preceding invocation phrase or word.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If an automated assistant requires invocation phrases or wake words before processing audio data, then the assistant can be reliably activated, but the interaction time is prolonged and resources are wasted during content rendering operations
Solution Approach 1:
The system applies different invocation requirements based on the local context of the operation being performed. During content rendering operations (TTS, alarm, notification), the system allows direct command execution without invocation phrases, while maintaining normal invocation requirements for other operations. This localized quality change resolves the contradiction by eliminating unnecessary delays in specific contexts while preserving reliable activation elsewhere.
Solution Approach 2:
The system dynamically adjusts its invocation behavior based on the current operational state. When detecting that a content rendering operation is active, the system transitions from requiring invocation phrases to accepting direct commands. This dynamic adaptation allows the system to respond to user needs in real-time without wasting resources on unnecessary wake word detection during relevant operations.
2Adaptability or versatility
If the automated assistant continuously detects invocation phrases during all operations, then the assistant remains responsive, but resource consumption increases unnecessarily during content rendering
Solution Approach 1:
Instead of continuously monitoring for invocation phrases during all operations, the system implements periodic or conditional monitoring. During content rendering operations, the system suspends wake word detection and focuses on processing direct commands related to the current content. This periodic action approach maintains response readiness for relevant operations while significantly reducing resource consumption during inappropriate monitoring periods.
Solution Approach 2:
The system determines its own operational state and automatically adjusts its monitoring behavior accordingly. When detecting that a content rendering operation is active, the system self-regulates by disabling unnecessary wake word detection and invocation phrase monitoring, thereby reducing resource consumption while maintaining responsiveness to relevant user commands.
3Ease of operation
If the automated assistant requires invocation phrases before executing operations, then the assistant maintains clear command separation, but the efficiency of task completion is reduced
Solution Approach 1:
The system applies different command processing rules based on the local operational context. During content rendering operations, the system accepts commands without invocation phrases, understanding that the context itself provides the necessary disambiguation. This localized quality change maintains command clarity in relevant contexts while dramatically improving task completion efficiency by eliminating unnecessary invocation steps.
Solution Approach 2:
The system uses feedback from the current operational state to determine the appropriate command processing mode. When detecting that a content rendering operation is active, the system adjusts its command parsing to accept direct commands without invocation phrases. This feedback mechanism ensures that command clarity is maintained through contextual understanding while improving efficiency by removing redundant invocation requirements.
Data Source
AI summary
Implementations relate to an automated assistant that is responsive, without requiring an invocation phrase or other invocation input(s), to certain spoken utterances when certain display content is being accessed by a user. The display content can be processed to identify certain inputs and/or other intents and parameters that are associated with assistant operations and are relevant to the display content. Thereafter, the automated assistant can determine whether any spoken utterances from the user correspond to those certain inputs, intents, and/or parameters. In response to receiving such a spoken utterance, the automated assistant can initialize performance of the relevant operation without necessitating that the user provides a preceding invocation phrase or other invocation input(s). When other display content is being accessed, the automated assistant can repeat the process for other inputs and operations.


