Dynamic Speech Recognition Model Update via Search Query Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems fail to accurately translate spoken words, especially when encountering new phrases or regional accents, as they rely on outdated language models that do not account for recent language usage or varied pronunciations.
Innovation Solution
A method and system for updating a speech recognition model by accessing a baseline model, obtaining information on recent language usage from search queries, and modifying sound occurrence probabilities, which includes generating a pronunciation dictionary using audio and transcript data to improve recognition accuracy across multiple users and devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a speech recognition system uses a static language model, then the system structure is simple and easy to implement, but the recognition accuracy deteriorates when encountering new words or regional accents
Solution Approach 1:
The patent implements dynamic language models that automatically update based on recent search queries and language usage data. The system transitions from static probability values to dynamically adjusted probabilities that adapt to new words, phrases, and pronunciation patterns without requiring manual retraining, thereby resolving the contradiction between simple structure and adaptability to new language usage
Solution Approach 2:
The system performs self-updating by automatically incorporating language patterns from aggregate search queries and user interactions. The language model autonomously adjusts probability distributions based on observed language usage, eliminating the need for external manual updates or individual user training sessions while maintaining recognition accuracy for evolving language patterns
2Measurement precision
If individual user training is implemented to accommodate regional accents, then recognition accuracy for that user improves, but the system complexity and training time increase
Solution Approach 1:
The patent creates a universal language model that serves all users simultaneously by incorporating diverse pronunciation patterns from aggregate search data. Instead of maintaining separate trained models for each user, the system universally applies probability adjustments derived from population-level language usage, accommodating regional accents across all users without individual training requirements
Solution Approach 2:
The system merges pronunciation data from multiple users and regional sources into a single unified language model. By combining diverse accent patterns from aggregate search queries and user interactions, the system creates a comprehensive model that recognizes various pronunciations without requiring separate training processes for each user
3Measurement precision
If the language model is updated frequently with recent language usage, then recognition accuracy for new words improves, but the computational resources and processing time increase
Solution Approach 1:
The patent applies partial updates to the language model by adjusting only the probability distributions for words and phrases that appear in recent search queries, rather than completely retraining the entire model. This selective updating approach incorporates new language usage with reduced computational overhead, balancing recognition accuracy for new words against resource consumption
Data Source
AI summary
A method for generating a speech recognition model includes accessing a baseline speech recognition model, obtaining information related to recent language usage from search queries, and modifying the speech recognition model to revise probabilities of a portion of a sound occurrence based on the information. The portion of a sound may include a word. Also, a method for generating a speech recognition model, includes receiving at a search engine from a remote device an audio recording and a transcript that substantially represents at least a portion of the audio recording, synchronizing the transcript with the audio recording, extracting one or more letters from the transcript and extracting the associated pronunciation of the one or more letters from the audio recording, and generating a dictionary entry in a pronunciation dictionary.


