Trigger Phrase Enrollment Quality Assessment for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems, particularly in always-on audio devices, face challenges in accurately recognizing trigger phrases due to background noise and poor enrollment quality, leading to degraded recognition accuracy.
Innovation Solution
An electronic device evaluates the quality of a spoken trigger phrase by measuring characteristics such as background noise level, length, and noise variability, and rejects the phrase if it does not meet predetermined thresholds, ensuring high-quality enrollment for improved recognition models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the device continuously listens for trigger phrases in always-on mode, then user convenience is improved by eliminating manual activation, but speech recognition accuracy deteriorates due to background noise and false triggers
Solution Approach 1:
The system performs preliminary voice activity detection and background noise assessment before full speech recognition processing. The device evaluates whether ambient noise levels are acceptable for accurate recognition, and only then activates the full speech recognition engine, preventing wasted processing on noisy segments
Solution Approach 2:
A noise assessment module acts as an intermediary between the always-on microphone and the speech recognition engine. This intermediary evaluates background noise characteristics and controls when the speech recognition engine should be activated, filtering out poor-quality audio segments before they reach the recognition system
2Reliability
If the trigger phrase recognizer is trained with multiple utterances, then recognition accuracy improves through model adaptation, but the enrollment time and complexity increase
Solution Approach 1:
The system collects more utterances than the minimum required for model training. By gathering excessive samples during enrollment, the system ensures sufficient data for robust model adaptation while maintaining a user-friendly process that doesn't feel overly lengthy to the end user
Solution Approach 2:
The enrollment process is designed to be autonomous and self-guided, with the system automatically managing recording, evaluation, and model training without requiring user intervention beyond providing the voice samples. The system self-assesses recording quality and requests re-recordings only when necessary
3Productivity
If the system accepts all recorded trigger phrases for training, then enrollment process simplifies and speeds up, but model quality deteriorates due to inclusion of noisy or poor-quality recordings
Solution Approach 1:
The system implements real-time feedback during the enrollment process by continuously assessing recording quality metrics such as signal-to-noise ratio, voice activity duration, and noise variability. Based on this feedback, the system immediately determines whether each recording should be accepted or rejected, allowing users to re-record problematic samples without delaying the overall process unnecessarily
Solution Approach 2:
The system dynamically adjusts evaluation thresholds and parameters based on the specific recording conditions. For example, it modifies the minimum voice activity duration requirement or noise level tolerance depending on the ambient environment detected during each recording attempt, optimizing the balance between acceptance rate and quality
4Adaptability or versatility
If the device processes speech recognition in noisy environments, then system versatility improves, but recognition accuracy deteriorates due to background noise interference
Solution Approach 1:
The system dynamically adjusts its operation mode based on real-time noise assessment. In low-noise environments, it performs full speech recognition processing, while in high-noise environments, it either requests the user to speak again or activates noise suppression algorithms, creating a dynamic response that adapts to changing acoustic conditions
Data Source
AI summary
An electronic device includes a microphone that receives an audio signal that includes a spoken trigger phrase, and a processor that is electrically coupled to the microphone. The processor measures characteristics of the audio signal, and determines, based on the measured characteristics, whether the spoken trigger phrase is acceptable for trigger phrase model training. If the spoken trigger phrase is determined not to be acceptable for trigger phrase model training, the processor rejects the trigger phrase for trigger phrase model training.


