User Device Activation Filtering via Audio Fingerprint Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
User devices are inadvertently activated by audio data containing words that resemble wakeup or activation words, either during the output of content assets or due to media mentions, leading to undesired wakeups.
Innovation Solution
A system that generates audio fingerprints and compares them to stored fingerprints associated with content assets to determine if a prospective wakeup word is actually a media mention or near-media mention, using block and allow lists to manage device activation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the user device activates upon receiving audio data containing words similar to wakeup words, then the device responds to user commands, but the device experiences undesired activations during content asset output
Solution Approach 1:
The system performs preliminary actions by generating audio fingerprints of content assets beforehand and storing them in a database. When audio data is received, the system compares the fingerprint of the incoming audio against the pre-stored fingerprints to determine if it matches content asset audio, thereby preventing undesired activations before they occur.
Solution Approach 2:
The system introduces an intermediary mechanism - audio fingerprint comparison - between the audio input and the activation decision. The fingerprint matching system acts as a mediator that filters out content asset audio from triggering activations, allowing only genuine user commands to activate the device.
2Measurement precision
If the device uses simple audio matching to detect wakeup words, then the activation process is fast, but the device cannot distinguish between actual wakeup words and media mentions
Solution Approach 1:
The system creates a copy of the audio data in the form of an audio fingerprint - a compact representation that captures essential acoustic features. This fingerprint copy is then compared against stored fingerprints, enabling accurate distinction between wakeup words and media mentions without requiring complex full-audio analysis.
Solution Approach 2:
The system transforms the audio data from its original complex waveform form into a fingerprint representation with specific parameters (spectral features, temporal patterns). This parameter transformation enables efficient comparison and accurate classification of audio content while maintaining computational feasibility.
Data Source
AI summary
Methods and systems are disclosed for determining a probability that a prospective wakeup or activation word is not an actual wakeup or activation word for a user device but instead is a word that has characteristics of, or otherwise sounds similar to, a wakeup or activation word and is received as a result of output of a content asset as opposed to being spoken by a user of the user device. Audio data associated with output of a content asset may be received and evaluated to determine if a prospective wakeup word in the audio data is an actual wakeup word or is, instead, not a wakeup word.


