Voice Device Wake Detection Using Directional Audio Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled audio devices struggle to differentiate between user-uttered and device-generated wake expressions, leading to unintended activation due to omnidirectional sound reflections and acoustic complexities in environments.

Innovation Solution

The audio device employs a microphone array with beamforming capabilities to generate directional audio signals, analyzing the number and pattern of these signals to determine if a wake expression is user-generated or device-generated, using machine learning techniques to learn and ignore self-generated expressions, and considering parameters like speaker output, echo characteristics, and loudness to make this distinction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the audio device uses omnidirectional microphone detection to capture wake expressions from all directions, then the device can detect user-uttered wake expressions from any position, but it also detects device-generated wake expressions due to sound reflections and acoustic echoes

Engineering Contradiction:
Improvewake expression detection coverageVSAvoidfalse activation rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the audio signal detection by creating multiple directional audio signals from the omnidirectional microphone input. Each directional signal focuses on a specific spatial sector, allowing the system to distinguish between sounds coming from different directions. This segmentation enables the device to identify whether a wake expression originates from the speaker direction (device-generated) or from other directions (user-uttered).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces directional audio signals as an intermediary layer between the omnidirectional microphone detection and the wake expression recognition. These directional signals act as mediators that process the raw audio input through spatial filtering, providing directional information that helps distinguish between user-uttered and device-generated wake expressions without requiring additional physical microphones.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the device speaks the wake expression aloud to provide feedback, then the user can confirm the device heard them, but the device may mistakenly detect its own spoken wake expression as a new user command

Engineering Contradiction:
Improveuser feedback confirmationVSAvoidunintended activation
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies preliminary action by analyzing the directional characteristics of the audio signal before triggering wake expression recognition. The system determines the direction of the detected wake expression and compares it with the speaker's orientation in advance. If the wake expression is detected coming from the speaker direction, the system preemptively prevents false activation by ignoring the detection or marking it as device-generated, thus avoiding the unintended activation problem before it occurs.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If the device uses simple wake expression detection without directional analysis, then the device complexity remains low, but the device cannot distinguish between user-uttered and device-generated wake expressions

Engineering Contradiction:
Improvesignal processing complexityVSAvoidwake expression origin identification
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent adds a spatial dimension to the wake expression detection by creating directional audio signals from omnidirectional microphone input. Instead of simply detecting the presence of a wake expression, the system analyzes the directional characteristics of the audio signal, effectively adding spatial information as another dimension to the detection process. This approach enables origin identification without requiring a complex array of multiple microphones, thus balancing measurement precision with device complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Effectively reduces false activations by accurately identifying user-uttered wake expressions, enhancing the reliability and precision of voice interaction systems in various environments.

Implementation Method 1

a microphone array that generates a plurality of directional audio signals

Methodology Applied
Scientific EffectAcoustic wave propagation: Sound

Implementation Method 2

The audio beamformer generates a plurality of directional audio signals based on the input audio

Methodology Applied
Scientific EffectBeamforming:

Implementation Method 3

omnidirectional sound reflections and acoustic complexities in environments

Methodology Applied
Scientific EffectAcoustic reflection: Reflection

Data Source

PatentUS11600271B2Detecting self-generated wake expressions
Publication Date: 2023.03.07 AMAZON TECH INC
  • US11600271B2 patent drawing
  • US11600271B2 patent drawing
  • US11600271B2 patent drawing

AI summary

A speech-based audio device may be configured to detect a user-uttered wake expression. For example, the audio device may generate a parameter indicating whether output audio is currently being produced by an audio speaker, whether the output audio contains speech, whether the output audio contains a predefined expression, loudness of the output audio, loudness of input audio, and/or an echo characteristic. Based on the parameter, the audio device may determine whether an occurrence of the predefined expression in the input audio is a result of an utterance of the predefined expression by a user.