Accent-Agnostic Wake Word Detection Using Unified Student Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated speech recognition (ASR) models require multiple wake word detection models to account for different languages and accents, leading to increased complexity and storage requirements.
Innovation Solution
A system and method for accent-agnostic frame-level wake word detection using a trained student model that is trained using audio samples in multiple accent types, allowing for a single model to recognize wake words across various accents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple wake word detection models are used to account for different languages and accents, then detection accuracy across diverse accents is improved, but device complexity and storage requirements increase
Solution Approach 1:
The patent combines multiple accent-specific wake word detection models into a single unified model. The system processes audio inputs through one model that has been trained to handle multiple accent types, eliminating the need for separate models for each accent while maintaining detection accuracy across diverse speech patterns
Solution Approach 2:
The unified wake word detection model is designed to perform multiple functions by detecting wake words across various accent types and languages simultaneously. This single model serves the role of what would otherwise require multiple specialized models, providing universal detection capability without requiring separate specialized components
2Reliability
If multiple wake word detection models are used to account for different languages and accents, then detection accuracy across diverse accents is improved, but storage requirements increase
Solution Approach 1:
The patent consolidates multiple accent-specific models into a single unified model, significantly reducing the storage space required. Instead of storing separate model parameters and weights for each accent type, the system stores one unified model that encompasses all accent variations, thereby reducing overall storage requirements while maintaining comprehensive detection capability
3Device complexity
If a single wake word detection model is used for multiple accents, then device complexity and storage requirements are reduced, but detection accuracy may deteriorate
Solution Approach 1:
The unified model undergoes preliminary training actions during the training phase where it is exposed to diverse accent data and learns to distinguish wake words across multiple accent types. This preliminary learning process enables the single model to achieve accuracy comparable to or better than multiple specialized models by pre-adapting to various speech patterns before actual deployment
Data Source
AI summary
A method includes accessing, using at least one processor of an electronic device, a machine learning model. The machine learning model is a trained student model that is trained using audio samples in a plurality of accent types. The method also includes receiving, using the at least one processor, an audio input from an audio input device. The method further includes providing, using the at least one processor, the audio input to the trained student model. The method also includes receiving, using the at least one processor, an output from the trained student model including frame-level probabilities associated with the audio input. In addition, the method includes instructing, using the at least one processor, at least one action based on the frame-level probabilities associated with the audio input.


