Wake Suppression for Audio Devices via Voice Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech-enabled devices often falsely trigger when they detect the wake word in audio streams from other devices, leading to unwanted activation and potential security breaches, especially when audio-playing devices play content that includes the wake word.

Innovation Solution

Implementing a system where the listening device uses a voice verification module to differentiate between live and reproduced audio, suppressing wake-up when the wake word is detected in audio streams from audio-playing devices, either through acoustic analysis or network communication, to prevent unnecessary activation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the listening device continuously monitors audio for wake words to enable responsive operation, then the device can quickly respond to user commands, but the device falsely triggers when detecting wake words in audio streams from other devices

Engineering Contradiction:
Improveresponse speedVSAvoidfalse trigger rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The listening device performs preliminary actions by analyzing audio characteristics before triggering wake-up. The system compares acoustic features such as spectral content, temporal patterns, and signal origin to determine whether the detected wake word comes from a nearby user or from an audio-playing device, thereby preventing false triggers while maintaining responsive operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary verification process between audio detection and wake-up triggering. This intermediary layer analyzes the audio signal characteristics and determines the source of the wake word before allowing state transition, effectively filtering out false triggers from audio streams while preserving legitimate user commands

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the device uses acoustic analysis to distinguish live and reproduced audio, then false triggers are reduced, but the device complexity increases

Engineering Contradiction:
Improvewake suppression accuracyVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio analysis process is segmented into multiple independent feature extraction modules, each analyzing specific acoustic characteristics such as spectral content, temporal patterns, and signal origin. This segmentation allows the system to handle complex audio differentiation through modular processing, reducing overall system complexity while maintaining high wake suppression accuracy

Inventive Principle:
Principle #1Segmentation

3Use of energy by stationary object

If the listening device remains in idle state to conserve energy, then power consumption is reduced, but the device cannot respond to wake words in audio streams

Engineering Contradiction:
Improvepower consumptionVSAvoidaudio stream awareness
Core Design Contradiction:
Use of energy by stationary objectVSAdaptability or versatility

Solution Approach 1:

The listening device performs partial audio analysis actions while remaining in idle state. Instead of fully processing all audio streams for wake words, the system performs selective monitoring that identifies audio streams from playing devices and suppresses wake-up triggers only for those streams, thereby maintaining energy efficiency while achieving appropriate adaptability

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11922939B2Wake suppression for audio playing and listening devices
Publication Date: 2024.03.05 SOUNDHOUND AI IP LLC
  • US11922939B2 patent drawing
  • US11922939B2 patent drawing
  • US11922939B2 patent drawing

AI summary

A system and method are disclosed for ignoring a wakeword received at a speech-enabled listening device when it is determined the wakeword is reproduced audio from an audio-playing device. Determination can be by detecting audio distortions, by an ignore flag sent locally between an audio-playing device and speech-enabled device, by and ignore flag sent from a server, by comparison of received audio played audio to a wakeword within an audio-playing device or a speech-enabled device, and other means.