Audio Watermarking for False Trigger Reduction in Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition systems often falsely trigger when detecting hotwords or keywords in audio streams, leading to unnecessary resource consumption, as they cannot distinguish between human utterances and pre-recorded or broadcasted content.

Innovation Solution

Incorporating an audio watermarking system where playback devices add ultrasonic signals to audio streams containing hotwords or keywords, allowing listening devices to differentiate between human inputs and pre-recorded content, thereby preventing false triggers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If speech recognition systems continuously monitor audio streams for hotwords, then responsiveness to user commands is improved, but false triggering from pre-recorded content increases

Engineering Contradiction:
Improveresponsiveness to user commandsVSAvoidfalse triggering rate
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent introduces an audio watermark as an intermediary signal that is embedded in pre-recorded content containing hotwords. This watermark acts as a mediator between the speech recognition system and the audio stream, providing additional information that helps the system distinguish between live user speech and pre-recorded content, thereby reducing false triggering while maintaining responsiveness

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter space by adding a new dimension for detection - the presence or absence of an audio watermark. Instead of relying solely on hotword detection, the system now considers both the hotword signal and the watermark signal, creating a multi-parameter detection approach that improves reliability without sacrificing speed

Inventive Principle:
Principle #35Parameter changes

2Reliability

If audio watermarking is added to distinguish pre-recorded content, then false triggering is reduced, but system complexity increases

Engineering Contradiction:
Improvefalse triggering rateVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex content analysis mechanisms with a simpler audio watermark detection approach. Instead of analyzing the entire audio stream to determine if it's pre-recorded, the system uses a dedicated watermark detector that looks for specific embedded signals, substituting a complex mechanical analysis system with a more straightforward detection mechanism

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The watermark is embedded in advance in the pre-recorded content, performing the classification work beforehand. This preliminary action allows the speech recognition system to make quick decisions during runtime without having to perform complex analysis, reducing the computational burden and system complexity during operation

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If hotwords in broadcasted content are detected and acted upon, then comprehensive speech recognition is achieved, but unnecessary resource consumption occurs

Engineering Contradiction:
Improvespeech recognition coverageVSAvoidcomputing resource consumption
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent extracts the watermark signal from the audio stream separately from the hotword detection process. By taking out this additional information channel, the system can make informed decisions about whether to process hotwords further, extracting only the necessary information (watermark presence) to avoid unnecessary resource consumption on false triggers

Inventive Principle:
Principle #2Taking out (Extraction)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution reduces the likelihood of false triggering, conserving computing resources by ensuring that actions are only taken in response to human inputs, thus optimizing resource usage.

Implementation Method 1

a playback device may analyze an audio stream for hotwords, keywords, or key phrases. Upon detection of a hotword, a keyword, or a key phrase, the playback device adds an audio watermark to the audio stream

Methodology Applied
Scientific EffectUltrasonic signal generation: Ultrasound

Data Source

PatentEP4202737B1Key phrase detection with audio watermarking
Publication Date: 2024.11.27 GOOGLE LLC
  • EP4202737B1 patent drawingFigure 1
  • EP4202737B1 patent drawingFigure 2
  • EP4202737B1 patent drawingFigure 3

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using audio watermarks with key phrases. One of the methods includes receiving, by a playback device, an audio data stream; determining, before the audio data stream is output by the playback device, whether a portion of the audio data stream encodes a particular key phrase by analyzing the portion using an automated speech recognizer; in response to determining that the portion of the audio data stream encodes the particular key phrase, modifying the audio data stream to include an audio watermark; and providing the modified audio data stream for output.