Audio Feature Normalization for Robust Wake Word Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal recognition systems face challenges in accurately detecting wake words and recognizing speech in varying acoustic conditions, often requiring extensive training data and being sensitive to background noise and microphone variations, which limits their robustness and reliability.
Innovation Solution
The method involves normalizing spectral features by determining a mean level-independent spectrum representation and performing a cepstral decomposition, followed by smoothing and subtracting it from feature vectors, to create a robust feature set that can be used for signal recognition processes, including wake word detection and speech recognition, even in new acoustic conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional audio signal recognition systems are used, then they can process audio data, but they are sensitive to background noise and microphone variations, reducing reliability
Solution Approach 1:
The patent applies parameter changes by normalizing spectral features through cepstral mean subtraction and variance normalization. This transforms the feature parameters to remove dependencies on acoustic conditions, microphone variations, and background noise levels, thereby improving reliability without requiring extensive retraining
Solution Approach 2:
The patent introduces cepstral coefficients as an intermediary representation that captures the essential spectral characteristics while being invariant to linear transformations in the acoustic environment. This intermediary feature representation mediates between the raw audio signal and the recognition model, filtering out harmful variations
2Measurement precision
If extensive training data is collected for different acoustic conditions, then recognition accuracy improves, but system complexity and data processing requirements increase
Solution Approach 1:
The patent extracts the essential acoustic characteristics through cepstral decomposition, separating the invariant speech features from the variable acoustic conditions. This extraction process creates a compact feature representation that captures recognition-critical information while eliminating the need for extensive training data across different acoustic environments
3Reliability
If feature normalization is applied to improve robustness, then false trigger rates reduce, but computational processing time increases
Solution Approach 1:
The patent performs preliminary normalization by pre-computing cepstral means and variances from training data, storing these as fixed transformation parameters. During runtime, the system applies these pre-computed parameters to normalize incoming features, significantly reducing real-time computational overhead while maintaining robustness against false triggers
Data Source
AI summary
A feature vector may be extracted from each frame of input digitized microphone audio data. The feature vector may include a power value for each frequency band of a plurality of frequency bands. A feature history data structure, including a plurality of feature vectors, may be formed. A normalized feature set that includes a normalized feature data structure may be produced by determining normalized power values for a plurality of frequency bands of each feature vector of the feature history data structure. A signal recognition or modification process may be based, at least in part, on the normalized feature data structure.


