ASR Language Model Generation via Search Engine Logs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic speech recognition (ASR) systems face challenges in scaling across different geo-linguistic locales and effectively capturing the relative importance of entities, particularly due to the computational burden of mining data from specific websites.

Innovation Solution

An improved ASR system generates language models based on expanded pattern data derived from search engine logs and a knowledge base, using a unified language model, time-sensitive language model, and class-based language model, which are trained on entity streams and pattern streams to adapt quickly to new applications and maintain minimal human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is mined from specific websites to generate language models, then language model generation is achieved, but computational burden increases and scalability across geo-linguistic locales becomes difficult

Engineering Contradiction:
Improvelanguage model generationVSAvoidcomputational burden
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces search engine logs as an intermediary data source between user queries and language model generation. Instead of directly mining websites, the system uses search logs that capture user interaction patterns, serving as a mediator that reduces computational complexity while maintaining language model quality across different geo-linguistic locales

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a simplified representation of website data through search engine logs. These logs copy essential information about user queries, selected entities, and search patterns without requiring full website scraping, thereby reducing computational burden while preserving the necessary language modeling data

Inventive Principle:
Principle #26Copying

2Reliability

If data is mined from specific websites to generate language models, then language model generation is achieved, but scalability across different geo-linguistic locales becomes hard

Engineering Contradiction:
Improvelanguage model generationVSAvoidscalability across geo-linguistic locales
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent makes search engine logs a universal data source that serves multiple geo-linguistic locales simultaneously. The same log processing pipeline can handle different languages and regions by filtering and processing logs according to locale-specific criteria, eliminating the need for separate mining processes for each locale

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies local quality by processing search logs with locale-specific filters and criteria. Each geo-linguistic locale receives customized processing based on its linguistic and cultural characteristics, allowing the system to adapt to local nuances while maintaining a unified overall architecture

Inventive Principle:
Principle #3Local quality

3Reliability

If data is mined from specific websites to generate language models, then language model generation is achieved, but relative importance of entities is not captured accurately

Engineering Contradiction:
Improvelanguage model generationVSAvoidentity importance capture
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback mechanisms through search engine logs that capture user interaction patterns. The logs record which entities users select, how long they spend on pages, and their search query patterns, providing feedback that enables accurate measurement of entity importance based on actual user behavior rather than static website content

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12087286B2Scalable entities and patterns mining pipeline to improve automatic speech recognition
Publication Date: 2024.09.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12087286B2 patent drawing
  • US12087286B2 patent drawing
  • US12087286B2 patent drawing

AI summary

A computing system obtains features that have been extracted from an acoustic signal, where the acoustic signal comprises spoken words uttered by a user. The computing system performs automatic speech recognition (ASR) based upon the features and a language model (LM) generated based upon expanded pattern data. The expanded pattern data includes a name of an entity and a search term, where the entity belongs to a segment identified in a knowledge base. The search term has been included in queries for entities belonging to the segment. The computing system identifies a sequence of words corresponding to the features based upon results of the ASR. The computing system transmits computer-readable text to a search engine, where the text includes the sequence of words.