Beam Group Selection and Merging for Noisy Speech Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing beam selection techniques in speech processing systems often fail to accurately identify desired audio signals under noisy conditions, leading to ineffective wakeword detection and speech processing performance, particularly when significant non-stationary noise is present, and beam switching occurs during speech utterances.
Innovation Solution
The system determines beam-specific signal quality metrics, including a minimum noise floor and signal-to-noise ratio (SNR), to select a group of beams based on low background noise and high SNR, and performs beam merging using a weighted sum to generate a combined output signal, prioritizing beams with favorable noise floor ratios and SNR values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional beam selection techniques are used, then the system can process audio signals, but it fails to accurately identify desired audio signals under noisy conditions
Solution Approach 1:
The patent changes the selection parameters from traditional single-beam approaches to multi-beam group selection based on noise floor ratios and signal-to-noise ratio (SNR) metrics. By evaluating multiple beams simultaneously and selecting groups with favorable noise characteristics, the system improves accuracy in noisy environments without being overwhelmed by interference noise.
2Adaptability or versatility
If beam switching occurs during speech utterances, then the system can adapt to different directions, but it causes instability in speech processing
Solution Approach 1:
The patent performs preliminary evaluation of multiple beams before final selection by calculating noise floor ratios and SNR metrics in advance. This preliminary assessment allows the system to pre-identify stable beam groups, reducing the likelihood of frequent switching during speech utterances and maintaining processing stability while still adapting to directional changes.
3Productivity
If the system processes audio from multiple directions, then it can capture more speech signals, but it increases complexity of beam selection
Solution Approach 1:
The patent merges multiple beam evaluation criteria into a unified selection process based on noise floor ratios and SNR metrics. Instead of processing beams independently and then selecting, the system combines the evaluation of multiple beams into a single multi-criteria assessment, capturing speech signals from multiple directions while managing complexity through integrated processing rather than separate stages.
Data Source
AI summary
A system that performs beam selection and beam merging using beam-specific signal quality metrics corresponding to a minimum noise floor. For example, a device may track a minimum noise floor for each beam, determine a highest minimum noise floor across the beams, and determine a noise floor ratio between the beam-specific minimum noise floor and the highest minimum noise floor. Using a combination of the noise floor ratio and signal-to-noise ratio (SNR) values, the device may perform beam selection by prioritizing low background noise as well as high SNR to select a pre-defined beam group. In addition, the device may use the noise floor ratio to perform beam merging and generate single-channel output audio data using the selected beam group. For example, the device may scale the beams based on a combination of the SNR value and the noise floor ratio.


