Voice Recognition Model Training via Noise Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems in devices like smartphones and tablets face challenges in accurately recognizing voice utterances, especially in varying audio environments due to background noise and individual speech characteristics such as accent and mispronunciations, requiring continuous training of phoneme or command databases.
Innovation Solution
The development of noise-based voice recognition model databases that can be manually or automatically trained using both live and previously-recorded utterances and noise samples, allowing for adaptive learning in different environments through directed and automated methods, including the combination of speech and noise signals for improved recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a phoneme or command database is trained to recognize individual speech characteristics, then recognition accuracy for that user improves, but the database becomes less accurate in different audio environments
Solution Approach 1:
The patent implements dynamic environmental noise classification that automatically detects and adapts to different audio environments (quiet, noisy, very noisy) in real-time. The system dynamically adjusts noise compensation strategies based on the detected environment, allowing the database to maintain accuracy across varying conditions rather than being fixed to a single training environment
Solution Approach 2:
The system changes operational parameters by applying different noise compensation techniques based on environmental classification. When noise is detected, the system modifies recognition parameters to account for noise characteristics, effectively adapting the database's behavior to match the current audio environment without requiring retraining
2Adaptability or versatility
If the voice recognition model database is continuously updated with new noise samples, then performance in diverse environments improves, but the complexity of the training process increases
Solution Approach 1:
The system performs automated environmental noise classification and self-trains using ambient noise captured during normal device operation. The device autonomously collects noise samples, classifies environmental types, and updates the voice recognition database without requiring manual user intervention or complex external training procedures
Solution Approach 2:
The system implements feedback loops where recognition results are continuously monitored and used to identify opportunities for database improvement. When speech recognition is performed in a new environment, the system feeds back the environmental noise characteristics to the training module, which automatically incorporates these into future recognition attempts
3Measurement precision
If manual training methods are used to update the voice recognition database, then training accuracy improves, but user time and effort increase
Solution Approach 1:
The system automatically performs environmental noise classification and database training without requiring manual user participation. The device autonomously monitors its audio environment, classifies noise types, and updates the voice recognition database in the background during normal operation, eliminating the need for users to spend time on manual training sessions
Data Source
AI summary
An electronic device digitally combines a single voice input with each of a series of noise samples. Each noise sample is taken from a different audio environment (e.g., street noise, babble, interior car noise). The voice input/noise sample combinations are used to train a voice recognition model database without the user having to repeat the voice input in each of the different environments. In one variation, the electronic device transmits the user's voice input to a server that maintains and trains the voice recognition model database.


