ASR Language Model Generation via Search Engine Logs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition (ASR) systems face challenges in scaling across different geo-linguistic locales and effectively capturing the relative importance of entities, particularly due to the computational burden of mining data from specific websites.
Innovation Solution
An improved ASR system generates language models based on expanded pattern data derived from search engine logs and a knowledge base, using a unified language model, time-sensitive language model, and class-based language model, which are trained on entity streams and pattern streams to adapt quickly to new applications and maintain minimal human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is mined from specific websites to generate language models, then language model generation is achieved, but computational burden increases and scalability across geo-linguistic locales becomes difficult
Solution Approach 1:
The patent introduces search engine logs as an intermediary data source between user queries and language model generation. Instead of directly mining websites, the system uses search logs that capture user interaction patterns, serving as a mediator that reduces computational complexity while maintaining language model quality across different geo-linguistic locales
Solution Approach 2:
The patent creates a simplified representation of website data through search engine logs. These logs copy essential information about user queries, selected entities, and search patterns without requiring full website scraping, thereby reducing computational burden while preserving the necessary language modeling data
2Reliability
If data is mined from specific websites to generate language models, then language model generation is achieved, but scalability across different geo-linguistic locales becomes hard
Solution Approach 1:
The patent makes search engine logs a universal data source that serves multiple geo-linguistic locales simultaneously. The same log processing pipeline can handle different languages and regions by filtering and processing logs according to locale-specific criteria, eliminating the need for separate mining processes for each locale
Solution Approach 2:
The patent applies local quality by processing search logs with locale-specific filters and criteria. Each geo-linguistic locale receives customized processing based on its linguistic and cultural characteristics, allowing the system to adapt to local nuances while maintaining a unified overall architecture
3Reliability
If data is mined from specific websites to generate language models, then language model generation is achieved, but relative importance of entities is not captured accurately
Solution Approach 1:
The patent incorporates feedback mechanisms through search engine logs that capture user interaction patterns. The logs record which entities users select, how long they spend on pages, and their search query patterns, providing feedback that enables accurate measurement of entity importance based on actual user behavior rather than static website content
Data Source
AI summary
A computing system obtains features that have been extracted from an acoustic signal, where the acoustic signal comprises spoken words uttered by a user. The computing system performs automatic speech recognition (ASR) based upon the features and a language model (LM) generated based upon expanded pattern data. The expanded pattern data includes a name of an entity and a search term, where the entity belongs to a segment identified in a knowledge base. The search term has been included in queries for entities belonging to the segment. The computing system identifies a sequence of words corresponding to the features based upon results of the ASR. The computing system transmits computer-readable text to a search engine, where the text includes the sequence of words.


