Audio Watermarking for False Trigger Reduction in Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition systems often falsely trigger when detecting hotwords or keywords in audio streams, leading to unnecessary resource consumption, as they cannot distinguish between human utterances and pre-recorded or broadcasted content.
Innovation Solution
Incorporating an audio watermarking system where playback devices add ultrasonic signals to audio streams containing hotwords or keywords, allowing listening devices to differentiate between human inputs and pre-recorded content, thereby preventing false triggers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If speech recognition systems continuously monitor audio streams for hotwords, then responsiveness to user commands is improved, but false triggering from pre-recorded content increases
Solution Approach 1:
The patent introduces an audio watermark as an intermediary signal that is embedded in pre-recorded content containing hotwords. This watermark acts as a mediator between the speech recognition system and the audio stream, providing additional information that helps the system distinguish between live user speech and pre-recorded content, thereby reducing false triggering while maintaining responsiveness
Solution Approach 2:
The system changes the parameter space by adding a new dimension for detection - the presence or absence of an audio watermark. Instead of relying solely on hotword detection, the system now considers both the hotword signal and the watermark signal, creating a multi-parameter detection approach that improves reliability without sacrificing speed
2Reliability
If audio watermarking is added to distinguish pre-recorded content, then false triggering is reduced, but system complexity increases
Solution Approach 1:
The patent replaces complex content analysis mechanisms with a simpler audio watermark detection approach. Instead of analyzing the entire audio stream to determine if it's pre-recorded, the system uses a dedicated watermark detector that looks for specific embedded signals, substituting a complex mechanical analysis system with a more straightforward detection mechanism
Solution Approach 2:
The watermark is embedded in advance in the pre-recorded content, performing the classification work beforehand. This preliminary action allows the speech recognition system to make quick decisions during runtime without having to perform complex analysis, reducing the computational burden and system complexity during operation
3Adaptability or versatility
If hotwords in broadcasted content are detected and acted upon, then comprehensive speech recognition is achieved, but unnecessary resource consumption occurs
Solution Approach 1:
The patent extracts the watermark signal from the audio stream separately from the hotword detection process. By taking out this additional information channel, the system can make informed decisions about whether to process hotwords further, extracting only the necessary information (watermark presence) to avoid unnecessary resource consumption on false triggers
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution reduces the likelihood of false triggering, conserving computing resources by ensuring that actions are only taken in response to human inputs, thus optimizing resource usage.
Implementation Method 1
a playback device may analyze an audio stream for hotwords, keywords, or key phrases. Upon detection of a hotword, a keyword, or a key phrase, the playback device adds an audio watermark to the audio stream
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using audio watermarks with key phrases. One of the methods includes receiving, by a playback device, an audio data stream; determining, before the audio data stream is output by the playback device, whether a portion of the audio data stream encodes a particular key phrase by analyzing the portion using an automated speech recognizer; in response to determining that the portion of the audio data stream encodes the particular key phrase, modifying the audio data stream to include an audio watermark; and providing the modified audio data stream for output.