Audio Feature Normalization for Robust Wake Word Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal recognition systems face challenges in accurately detecting wake words and recognizing speech in varying acoustic conditions, often requiring extensive training data and being sensitive to background noise and microphone variations, which limits their robustness and reliability.

Innovation Solution

The method involves normalizing spectral features by determining a mean level-independent spectrum representation and performing a cepstral decomposition, followed by smoothing and subtracting it from feature vectors, to create a robust feature set that can be used for signal recognition processes, including wake word detection and speech recognition, even in new acoustic conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional audio signal recognition systems are used, then they can process audio data, but they are sensitive to background noise and microphone variations, reducing reliability

Engineering Contradiction:
Improverobustness of signal recognitionVSAvoidsensitivity to background noise and acoustic conditions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies parameter changes by normalizing spectral features through cepstral mean subtraction and variance normalization. This transforms the feature parameters to remove dependencies on acoustic conditions, microphone variations, and background noise levels, thereby improving reliability without requiring extensive retraining

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces cepstral coefficients as an intermediary representation that captures the essential spectral characteristics while being invariant to linear transformations in the acoustic environment. This intermediary feature representation mediates between the raw audio signal and the recognition model, filtering out harmful variations

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive training data is collected for different acoustic conditions, then recognition accuracy improves, but system complexity and data processing requirements increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidtraining data requirements and processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential acoustic characteristics through cepstral decomposition, separating the invariant speech features from the variable acoustic conditions. This extraction process creates a compact feature representation that captures recognition-critical information while eliminating the need for extensive training data across different acoustic environments

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If feature normalization is applied to improve robustness, then false trigger rates reduce, but computational processing time increases

Engineering Contradiction:
Improvefalse trigger rateVSAvoidfeature processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary normalization by pre-computing cepstral means and variances from training data, storing these as fixed transformation parameters. During runtime, the system applies these pre-computed parameters to normalize incoming features, significantly reducing real-time computational overhead while maintaining robustness against false triggers

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12175965B2Method and apparatus for normalizing features extracted from audio data for signal recognition or modification
Publication Date: 2024.12.24 DOLBY LABORATORIES LICENSING CORP
  • US12175965B2 patent drawing
  • US12175965B2 patent drawing
  • US12175965B2 patent drawing

AI summary

A feature vector may be extracted from each frame of input digitized microphone audio data. The feature vector may include a power value for each frequency band of a plurality of frequency bands. A feature history data structure, including a plurality of feature vectors, may be formed. A normalized feature set that includes a normalized feature data structure may be produced by determining normalized power values for a plurality of frequency bands of each feature vector of the feature history data structure. A signal recognition or modification process may be based, at least in part, on the normalized feature data structure.