ASR Model Switching via GPS and Sensor Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition (ASR) systems face challenges in varying noise environments and types of speech, leading to suboptimal performance due to static tuning parameters and inability to independently minimize false accepts and false rejects across different environments.
Innovation Solution
The ASR system collects geographic location data and sensor data to dynamically adjust acoustic models and speech recognition models based on identified action profiles, compensating for changes in noise environments and user speech patterns, allowing for independent optimization of trigger and command speech recognizers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If ASR algorithms use static tuning parameters trained for a specific noise environment, then the system works well in that expected environment, but the ASR system fails to properly recognize speech when the environment deviates from training conditions
Solution Approach 1:
The patent implements dynamic tuning parameters that automatically adjust based on detected noise characteristics. The system transitions from static parameters to dynamic parameters that adapt in real-time to changing acoustic environments, allowing the ASR system to maintain high recognition accuracy across diverse conditions without requiring retraining for each environment.
Solution Approach 2:
The system changes acoustic parameters dynamically by analyzing noise characteristics and adjusting tuning parameters accordingly. This includes modifying parameters such as noise suppression levels, acoustic models, and recognition thresholds based on the detected environment, enabling the ASR system to optimize performance for each specific acoustic condition.
2Measurement precision
If ASR algorithms are tuned to minimize false accepts in quiet environments, then recognition precision improves in quiet settings, but the system produces more false rejects in noisier environments
Solution Approach 1:
The system dynamically adjusts recognition thresholds and tuning parameters based on detected noise levels. In quiet environments, parameters are optimized for high precision with stricter matching criteria, while in noisy environments, parameters shift to prioritize reliability with more tolerant matching, thereby reducing false rejects without sacrificing too much precision.
Solution Approach 2:
The patent implements dynamic threshold adjustment where the decision boundaries for speech recognition are not fixed but adapt based on the acoustic environment. This dynamic approach allows the system to balance precision and reliability in real-time, switching between conservative and liberal recognition strategies depending on noise conditions.
3Reliability
If ASR algorithms are tuned to minimize false rejects in noisy environments, then recognition reliability improves in noisy settings, but the system produces more false accepts in quieter environments
Solution Approach 1:
The system dynamically modifies recognition parameters based on environmental noise characteristics. When noise is detected, parameters are adjusted to favor reliability with lower thresholds and more permissive matching. When the environment is quiet, parameters shift back to prioritize precision with higher thresholds and stricter verification, thereby minimizing false accepts.
4Device complexity
If a single ASR algorithm is used for both trigger speech recognition and command speech recognition, then system complexity is reduced, but false accepts and false rejects cannot be independently minimized for each recognizer
Solution Approach 1:
The patent divides the speech recognition system into separate trigger and command recognizers, each with independent tuning parameters and optimization goals. This segmentation allows the trigger recognizer to be optimized for wake-word detection with high reliability, while the command recognizer can be optimized for instruction recognition with high precision, without the performance of one compromising the other.
Data Source
AI summary
An automatic speech recognition (ASR) system is disclosed that compensates for different noise environments and types of speech. The ASR system may be implemented as part of an action camera that collects status data, such as geographic location data and/or sensor data. The ASR system may perform speech recognition using an acoustic model and a speech recognition model, which are trained for operation in specific noise environments and/or for specific types of speech. The computing device may categorize a current status of the action camera, as indicated by the status data, into an action profile, which may represent a particular activity (e.g., running, cycling, etc.) or state of the computing device. The computing device may dynamically switch the acoustic model and/or the speech recognition model to compensate for anticipated changes in the noise environment and speech based upon the action profile to facilitate the recognition of various action camera functions.


