Beamformer Trigger Word Detection for Audio Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing systems face challenges with bandwidth consumption, privacy concerns, and inefficient resource utilization due to continuous audio transmission and ineffective beam selection methods, particularly in noisy environments, which affect wakeword detection and overall speech recognition performance.
Innovation Solution
A device employing a low-power beam-based trigger word detection system that uses multiple trained models to select the most relevant audio beam for further processing, based on confidence scores, and incorporates an adaptive beamformer to isolate and enhance desired audio signals while reducing noise, thereby optimizing resource usage and improving wakeword detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If continuous audio transmission is used for speech processing, then speech recognition can be performed, but bandwidth consumption increases and privacy concerns arise
Solution Approach 1:
The system performs preliminary beam selection and trigger word detection locally on the device before transmitting audio data. Multiple beams are pre-processed to identify potential trigger words, and only relevant audio segments are transmitted for full speech processing, reducing bandwidth consumption while maintaining recognition accuracy
Solution Approach 2:
The invention extracts and transmits only the most relevant audio beam containing the trigger word rather than continuous audio streams. By selecting specific beams with high probability of containing trigger words and transmitting only those segments, the system reduces bandwidth usage while preserving speech recognition capability
2Measurement precision
If multiple audio beams are processed continuously, then wakeword detection accuracy can be improved, but computational resources are inefficiently utilized
Solution Approach 1:
The system applies partial processing to multiple beams by computing trigger word probabilities for each beam and selectively forwarding only high-probability beams for complete speech processing. This partial action approach maintains detection accuracy by considering multiple beams while improving resource efficiency by avoiding full processing of all beams
Solution Approach 2:
The audio processing is segmented into multiple independent beams, each processed through a lightweight trigger word detection model. This segmentation allows parallel processing of multiple beams with reduced computational overhead per beam, improving both accuracy through multi-beam analysis and efficiency through divided computational workload
3Loss of information
If audio from all directions is transmitted, then no audio information is lost, but bandwidth is wasted and privacy is compromised
Solution Approach 1:
The system applies different processing qualities to different audio beams based on their relevance. High-probability beams containing trigger words are processed with full quality and transmitted, while low-probability beams are either processed locally without transmission or processed with reduced quality, optimizing bandwidth usage while preserving critical audio information
Data Source
AI summary
An audio capture device that incorporates a beamformer and beam-specific trigger word detection. Audio data from each beam is processed by a low power trigger word detector, such as a neural network or other trained model to detect if audio data (such as an audio frame or feature vector corresponding thereto) likely includes part of a trigger word. The beam that either most strongly represents a trigger word portion or represents a trigger word portion most early in time may be selected for further processing such as speech processing or confirmation by a more robust power intensive trigger word detector.


