Beamformer Trigger Word Detection for Audio Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems face challenges with bandwidth consumption, privacy concerns, and inefficient resource utilization due to continuous audio transmission and ineffective beam selection methods, particularly in noisy environments, which affect wakeword detection and overall speech recognition performance.

Innovation Solution

A device employing a low-power beam-based trigger word detection system that uses multiple trained models to select the most relevant audio beam for further processing, based on confidence scores, and incorporates an adaptive beamformer to isolate and enhance desired audio signals while reducing noise, thereby optimizing resource usage and improving wakeword detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If continuous audio transmission is used for speech processing, then speech recognition can be performed, but bandwidth consumption increases and privacy concerns arise

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system performs preliminary beam selection and trigger word detection locally on the device before transmitting audio data. Multiple beams are pre-processed to identify potential trigger words, and only relevant audio segments are transmitted for full speech processing, reducing bandwidth consumption while maintaining recognition accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention extracts and transmits only the most relevant audio beam containing the trigger word rather than continuous audio streams. By selecting specific beams with high probability of containing trigger words and transmitting only those segments, the system reduces bandwidth usage while preserving speech recognition capability

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If multiple audio beams are processed continuously, then wakeword detection accuracy can be improved, but computational resources are inefficiently utilized

Engineering Contradiction:
Improvewakeword detection accuracyVSAvoidresource utilization efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system applies partial processing to multiple beams by computing trigger word probabilities for each beam and selectively forwarding only high-probability beams for complete speech processing. This partial action approach maintains detection accuracy by considering multiple beams while improving resource efficiency by avoiding full processing of all beams

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The audio processing is segmented into multiple independent beams, each processed through a lightweight trigger word detection model. This segmentation allows parallel processing of multiple beams with reduced computational overhead per beam, improving both accuracy through multi-beam analysis and efficiency through divided computational workload

Inventive Principle:
Principle #1Segmentation

3Loss of information

If audio from all directions is transmitted, then no audio information is lost, but bandwidth is wasted and privacy is compromised

Engineering Contradiction:
Improveaudio information completenessVSAvoidbandwidth usage
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The system applies different processing qualities to different audio beams based on their relevance. High-probability beams containing trigger words are processed with full quality and transmitted, while low-probability beams are either processed locally without transmission or processed with reduced quality, optimizing bandwidth usage while preserving critical audio information

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10304475B1Trigger word based beam selection
Publication Date: 2019.05.28 AMAZON TECH INC
  • US10304475B1 patent drawing
  • US10304475B1 patent drawing
  • US10304475B1 patent drawing

AI summary

An audio capture device that incorporates a beamformer and beam-specific trigger word detection. Audio data from each beam is processed by a low power trigger word detector, such as a neural network or other trained model to detect if audio data (such as an audio frame or feature vector corresponding thereto) likely includes part of a trigger word. The beam that either most strongly represents a trigger word portion or represents a trigger word portion most early in time may be selected for further processing such as speech processing or confirmation by a more robust power intensive trigger word detector.