Voice Recognition Model Training via Noise Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems in devices like smartphones and tablets face challenges in accurately recognizing voice utterances, especially in varying audio environments due to background noise and individual speech characteristics such as accent and mispronunciations, requiring continuous training of phoneme or command databases.

Innovation Solution

The development of noise-based voice recognition model databases that can be manually or automatically trained using both live and previously-recorded utterances and noise samples, allowing for adaptive learning in different environments through directed and automated methods, including the combination of speech and noise signals for improved recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a phoneme or command database is trained to recognize individual speech characteristics, then recognition accuracy for that user improves, but the database becomes less accurate in different audio environments

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic environmental noise classification that automatically detects and adapts to different audio environments (quiet, noisy, very noisy) in real-time. The system dynamically adjusts noise compensation strategies based on the detected environment, allowing the database to maintain accuracy across varying conditions rather than being fixed to a single training environment

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters by applying different noise compensation techniques based on environmental classification. When noise is detected, the system modifies recognition parameters to account for noise characteristics, effectively adapting the database's behavior to match the current audio environment without requiring retraining

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the voice recognition model database is continuously updated with new noise samples, then performance in diverse environments improves, but the complexity of the training process increases

Engineering Contradiction:
Improveenvironmental coverageVSAvoidtraining process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs automated environmental noise classification and self-trains using ambient noise captured during normal device operation. The device autonomously collects noise samples, classifies environmental types, and updates the voice recognition database without requiring manual user intervention or complex external training procedures

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback loops where recognition results are continuously monitored and used to identify opportunities for database improvement. When speech recognition is performed in a new environment, the system feeds back the environmental noise characteristics to the training module, which automatically incorporates these into future recognition attempts

Inventive Principle:
Principle #23Feedback

3Measurement precision

If manual training methods are used to update the voice recognition database, then training accuracy improves, but user time and effort increase

Engineering Contradiction:
Improvetraining accuracyVSAvoiduser training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically performs environmental noise classification and database training without requiring manual user participation. The device autonomously monitors its audio environment, classifies noise types, and updates the voice recognition database in the background during normal operation, eliminating the need for users to spend time on manual training sessions

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9275638B2Method and apparatus for training a voice recognition model database
Publication Date: 2016.03.01 GOOGLE TECHNOLOGY HOLDINGS LLC
  • US9275638B2 patent drawing
  • US9275638B2 patent drawing
  • US9275638B2 patent drawing

AI summary

An electronic device digitally combines a single voice input with each of a series of noise samples. Each noise sample is taken from a different audio environment (e.g., street noise, babble, interior car noise). The voice input/noise sample combinations are used to train a voice recognition model database without the user having to repeat the voice input in each of the different environments. In one variation, the electronic device transmits the user's voice input to a server that maintains and trains the voice recognition model database.