Virtual Assistant Wake-Up Phrase Configuration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual assistant systems face inefficiencies in language recognition and power consumption due to large vocabulary sizes and the need for continuous listening, particularly when recognizing multiple languages, which affects accuracy and efficiency.
Innovation Solution
The system configures behavior based on distinct wake-up phrases to select appropriate language vocabularies, text-to-speech systems, acoustic models, and other attributes, allowing for efficient processing of utterances by switching between different language models and features based on detected wake-up phrases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a joint recognizer combining vocabularies of multiple languages is used, then the system can recognize speech in multiple languages, but the vocabulary size becomes too large and recognition accuracy decreases
Solution Approach 1:
The patent divides the large multilingual vocabulary into separate language-specific vocabularies. Instead of using one joint recognizer with a combined vocabulary of multiple languages, the system creates multiple independent recognizers, each specialized for a single language. This segmentation reduces the vocabulary size for each recognizer, thereby improving recognition accuracy while maintaining multi-language capability through selective activation based on wake-up phrase detection.
2Ease of operation
If continuous listening with large-vocabulary ASR is performed, then the system can recognize sophisticated utterances, but power consumption becomes impractical for mobile devices
Solution Approach 1:
The patent implements a dynamic listening mode that switches between two states: a low-power phrase-spotting state that continuously monitors for wake-up phrases, and a high-power full ASR state that activates only when a wake-up phrase is detected. This dynamic transition allows the system to maintain sophisticated utterance processing capability when needed while dramatically reducing power consumption during normal operation, making it practical for mobile devices.
3Use of energy by moving object
If phrase spotting is used to detect wake-up phrases, then power consumption is reduced, but the system may behave intrusively by responding to utterances not intended for it
Solution Approach 1:
The patent introduces distinct wake-up phrases as intermediaries between the user and the virtual assistant. Instead of continuously listening and potentially misinterpreting random utterances, the system requires a specific wake-up phrase to activate full ASR processing. This intermediary mechanism ensures that the system only responds when explicitly invoked, eliminating intrusive behavior while maintaining low power consumption during the phrase-spotting state.
Data Source
AI summary
A speech-enabled dialog system responds to a plurality of wake-up phrases. Based on which wake-up phrase is detected, the system's configuration is modified accordingly. Various configurable aspects of the system include selection and morphine of a text-to-speech voice; configuration of acoustic model, language model, vocabulary, and grammar; configuration of a graphic animation; configuration of virtual assistant personality parameters; invocation of a particular user profile; invocation of an authentication function; and configuration of an open sound. Configuration depends on a target market segment. Configuration also depends on the state of the dialog system, such as whether a previous utterance was an information query.


