Multi-factor Audio Watermarking for Automated Assistant Activation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated assistants often experience unintended activations when hotwords are detected in audio from media content, leading to negative user experiences, resource wastage, and the need for users to undo undesired actions.

Innovation Solution

Implementing a multi-factor audio watermarking system that combines imperceptible audio watermarks with additional factors like speech-based signals to detect media content, thereby suppressing processing of queries included in media audio and adapting hotword detection and speech recognition accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If hotword detection is used to activate automated assistant functions, then the automated assistant can respond to user requests, but unintended activations occur when hotwords are present in media content

Engineering Contradiction:
Improveautomated assistant activationVSAvoidactivation accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

An audio watermark detection system serves as an intermediary between hotword detection and automated assistant activation. The watermark detector analyzes audio content to identify media content, and when media content is detected, it suppresses automated assistant activation even if hotwords are present. This intermediary mechanism resolves the contradiction by adding a verification layer that distinguishes between legitimate user requests and hotwords embedded in media content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If audio processing is continuously performed to detect hotwords, then the automated assistant can respond quickly to user requests, but computational resources are wasted when no user input is intended

Engineering Contradiction:
Improveresponse timeVSAvoidcomputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts its processing mode based on watermark detection results. When media content is detected through watermark analysis, the system transitions to a low-power mode where full audio processing and automated assistant functions are suppressed. When no media content is detected, the system operates in normal mode with full processing capabilities. This dynamic adaptation resolves the contradiction by optimizing resource consumption while maintaining quick response capability when needed.

Inventive Principle:
Principle #15Dynamics

3Productivity

If hotword detection threshold is lowered to reduce false negatives, then more user requests are captured, but false positives increase leading to more unintended activations

Engineering Contradiction:
Improverequest capture rateVSAvoidactivation accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The audio watermark detection system acts as a mediator that validates hotword detections before triggering automated assistant activation. Even when the hotword detection threshold is lowered to capture more user requests, the watermark detection layer provides an additional verification step that filters out false positives from media content. This resolves the contradiction by allowing aggressive hotword detection parameters while maintaining high activation accuracy through the intermediary validation mechanism.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12254888B2Multi-factor audio watermarking
Publication Date: 2025.03.18 GOOGLE LLC
  • US12254888B2 patent drawing
  • US12254888B2 patent drawing
  • US12254888B2 patent drawing

AI summary

Techniques are described herein for multi-factor audio watermarking. A method includes: receiving audio data; processing the audio data to generate predicted output that indicates a probability of one or more hotwords being present in the audio data; determining that the predicted output satisfies a threshold that is indicative of the one or more hotwords being present in the audio data; in response to determining that the predicted output satisfies the threshold, processing the audio data using automatic speech recognition to generate a speech transcription feature; detecting a watermark that is embedded in the audio data; and in response to detecting the watermark: determining that the speech transcription feature corresponds to one of a plurality of stored speech transcription features; and in response to determining that the speech transcription feature corresponds to one of the plurality of stored speech transcription features, suppressing processing of a query included in the audio data.