Virtual Assistant Wake-Up Filter for False Activation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional virtual assistant applications often activate unnecessarily due to background noise, leading to distractions and potential hazards, such as car accidents, as they fail to differentiate between human-generated and electronically reproduced wake-up words.
Innovation Solution
Incorporating a signal augmentation pattern with 'gaps' in the audio signal, processed to exclude specific frequency levels, allowing the virtual assistant to recognize and ignore electronically reproduced wake-up words, thereby preventing false activations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the virtual assistant application uses conventional audio recognition to detect wake-up words, then it can respond to user commands, but it suffers from false activations caused by background noise from electronic devices
Solution Approach 1:
The patent introduces an intermediary mechanism - a spectral fingerprint analysis system - that acts as a mediator between the audio input and the wake-up word recognition. This intermediary analyzes the spectral characteristics of the audio signal and compares it against known patterns from electronic devices, enabling the system to distinguish between human speech and electronically reproduced audio without affecting the normal wake-up word functionality
Solution Approach 2:
The patent changes the parameter space by moving from simple audio threshold detection to spectral fingerprint analysis. By transforming the audio signal into the frequency domain and analyzing spectral characteristics, the system can identify patterns unique to electronic device audio output, thereby resolving the contradiction between maintaining wake-up sensitivity and rejecting false activations
2Object-affected harmful factors
If the virtual assistant application filters out all background audio signals, then false activations are reduced, but legitimate user wake-up words may also be missed
Solution Approach 1:
The patent applies local quality by analyzing specific spectral regions and characteristics rather than uniformly filtering all audio signals. By examining the local spectral fingerprint - specific frequency patterns, harmonic structures, and spectral envelope characteristics - the system can selectively identify electronic device audio while preserving human speech detection in other spectral regions
3Object-affected harmful factors
If the system implements sophisticated audio analysis to distinguish human speech from electronic audio, then false activations are reduced, but system complexity increases
Solution Approach 1:
The patent uses copying by creating spectral fingerprint representations of electronic device audio outputs. Instead of implementing complex real-time analysis of all audio characteristics, the system captures and stores spectral patterns from known electronic device wake-up word reproductions, then compares incoming audio against these copied patterns, significantly reducing processing complexity while maintaining accuracy
Data Source
AI summary
There is provided a method for operating a speaker device able to be activated by receiving and recognizing a predetermined wake up word. The method is executable at a server. The method comprises: capturing, by the speaker device, an audio signal having been generated in a vicinity of the speaker device; retrieving, by the speaker device, a processing filter, the processing filter being indicative of a pre-determined signal augmentation pattern representative of an excluded portion that has been excluded from an originating utterance having the wake up word, the originating utterance to be reproduced by an other electronic device; applying, by the speaker device, the processing filter to determine presence of the pre-determined signal augmentation pattern in the audio signal; based on determining the presence of the pre-determined signal augmentation pattern in the audio signal, determining that the audio signal has been produced by the other electronic device.


