Phase Inversion for Wake Word Detection in Audio Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices with hardware configurations where speakers are close to microphones struggle to detect wake words or voice commands due to interference from playing audio content, requiring users to speak louder or move closer, which is inconvenient.

Innovation Solution

The phase inversion feature reduces latency and noise interference by inverting and merging audio waveforms from the source content and user input, allowing for more accurate detection of wake words or voice commands even when the device is playing audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Volume of moving object

If the speaker is placed close to the microphone for compact device design, then the device size is reduced, but the wake word detection accuracy deteriorates due to audio interference

Engineering Contradiction:
Improvedevice sizeVSAvoidwake word detection accuracy
Core Design Contradiction:
Volume of moving objectVSMeasurement precision

Solution Approach 1:

The patent converts the harmful audio interference from the speaker into a beneficial signal by capturing it with the microphone, then using phase inversion to cancel it out. The interference that initially degrades wake word detection is transformed into a tool for improving detection accuracy by subtracting the known speaker signal from the microphone input.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The system performs preliminary anti-action by inverting the phase of the speaker's audio output before combining it with the microphone signal. This pre-inversion creates destructive interference with the original speaker signal, effectively canceling it out before it can interfere with wake word detection.

Inventive Principle:
Principle #9Preliminary anti-action

2Adaptability or versatility

If the device plays audio content through the speaker, then the entertainment functionality is provided, but the voice command detection reliability deteriorates due to noise interference

Engineering Contradiction:
Improveentertainment functionalityVSAvoidvoice command detection reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the audio signal into two distinct components: the speaker output signal and the microphone captured signal. By separating these signals and processing them independently (inverting one and combining with the other), the system can eliminate the speaker interference while preserving the voice command information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processed audio signal acts as an intermediary between the speaker and the wake word detection system. By inserting this intermediate signal processing step (phase inversion and combination), the system mediates the interaction between speaker output and microphone input, allowing both functions to coexist without interference.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If the user speaks louder or moves closer to compensate for audio interference, then the wake word detection reliability improves, but the ease of operation deteriorates

Engineering Contradiction:
Improvewake word detection reliabilityVSAvoiduser convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs self-service by automatically compensating for the audio interference through signal processing. Instead of requiring the user to adjust their behavior (speak louder or move closer), the device independently processes its own signals to eliminate interference, making the system self-correcting and user-friendly.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient identification of wake words in real-time without requiring users to adjust their volume or distance, allowing for seamless voice command execution without delay.

Implementation Method 1

The phase inversion module may be configured to receive one or more variable waveforms, identify latency between the one or more variable waveforms, reduce the latency, invert the phase of a modified variable waveform, merge the modified variable waveform with the other variable waveforms to generate a merged variable waveform

Methodology Applied
Scientific EffectPhase inversion:

Data Source

PatentUS10699729B1Phase inversion for virtual assistants and mobile music apps
Publication Date: 2020.06.30 AMAZON TECH INC
  • US10699729B1 patent drawing
  • US10699729B1 patent drawing
  • US10699729B1 patent drawing

AI summary

Techniques for identifying a wake word by a device that is also playing audio content at the same time are described herein. For example, a device may execute playback of an audio file with a corresponding first variable wave form. The device may receive a second variable wave form that includes the first variable wave form and additional audio. In embodiments, a latency value may be identified based on comparing amplitudes and frequencies of portions of the first variable wave form and the second variable wave form. The second variable wave form may be modified by applying the latency value and inverting the second variable wave form with respect to the first variable wave form. The modified variable wave form may be merged with the first variable wave form to generate a merged variable wave form. A particular audio signal may be identified in the merged variable wave form.