Web Data Language Model Generation for Spoken Dialog Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Creating statistical language models for voice-enabled automated call center applications is hindered by the difficulty in collecting sufficient data, as web language statistics differ significantly from conversational styles, leading to resource-intensive and time-consuming model training processes.

Innovation Solution

A system and method that filters web data to remove unwanted information, extracts predicate/argument pairs, generates conversational utterances by merging these pairs into templates, and creates a language model using the filtered data, addressing the disparity between web and conversational language statistics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If web data is used to train language models, then the quantity of training data is improved, but the quality and applicability to conversational speech deteriorates due to statistical differences between web language and conversational style

Engineering Contradiction:
Improvequantity of training dataVSAvoidquality and applicability of language model
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent extracts useful linguistic patterns, vocabulary, and domain-specific information from web data while discarding irrelevant web-specific sequences and disfluencies. This selective extraction allows the system to leverage the quantity advantage of web data while maintaining quality standards appropriate for conversational speech recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing layer that transforms raw web data into a form suitable for training conversational language models. This intermediary process includes filtering, normalization, and adaptation steps that bridge the statistical gap between web language and conversational speech, enabling effective knowledge transfer from the web domain.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If traditional data collection methods are used for voice-enabled applications, then the quality and relevance of training data is improved, but the time and resources required for data collection and model training increase significantly

Engineering Contradiction:
Improvequality and relevance of training dataVSAvoidtime-to-deployment
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of web data in advance, creating pre-processed training datasets that can be directly used for model training. This preliminary action includes filtering, extracting relevant patterns, and preparing the data in a format ready for training, thereby eliminating the need for time-consuming data collection and preprocessing steps during deployment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic conversational utterances by merging extracted predicate/argument pairs into conversational templates, generating artificial training data that mimics real conversational speech. This copying approach allows the system to leverage abundant web data to create quality conversational training samples without requiring actual recorded speech collections.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If web-specific word sequences are included in the language model, then the coverage of domain-specific vocabulary is improved, but the naturalness and fluency of spoken language generation deteriorates

Engineering Contradiction:
Improvecoverage of domain-specific vocabularyVSAvoidnaturalness and fluency of speech
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent applies different processing rules to different parts of the web data. Domain-specific vocabulary, product names, and key phrases are preserved and emphasized, while web-specific sequences and unnatural constructions are filtered out. This local quality approach ensures that useful domain terminology is maintained while natural speech patterns are preserved in the generated language model.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9299345B1Bootstrapping language models for spoken dialog systems using the world wide web
Publication Date: 2016.03.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9299345B1 patent drawing
  • US9299345B1 patent drawing
  • US9299345B1 patent drawing

AI summary

A system, method and computer readable medium that generates a language model from data from a web domain is disclosed. The method may include filtering web data to remove unwanted data from the web domain data, extracting predicate/argument pairs from the filtered web data, generating conversational utterances by merging the extracted predicate/argument pairs into conversational templates, and generating a web data language model using the generated conversational utterances.