Modified Language Model for Disfluency Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face limitations in handling disfluencies, such as filled pauses, repetitions, and repairs, which lead to inefficiencies and errors in free form, continuous speech recognition due to their inability to account for these speech patterns in language models.

Innovation Solution

A modified language model is introduced that includes synthetic n-grams, which incorporate disfluency tokens to improve the recognition of speech by adapting the existing clean language model to account for disfluencies, allowing for more accurate matching of phoneme sequences and reducing errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a clean language model is used for speech recognition, then the system is simpler and faster to deploy, but it cannot handle disfluencies in free form speech

Engineering Contradiction:
Improveability to handle disfluenciesVSAvoidlanguage model complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-generating synthetic n-grams that incorporate disfluency tokens and integrating them into the language model before deployment. This allows the system to handle disfluencies without requiring complex real-time processing during speech recognition, as the disfluency patterns are already accounted for in the pre-built language model structure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces disfluency tokens as intermediary elements that bridge clean speech patterns and actual disfluent speech. These tokens act as mediators within the language model, allowing the system to recognize and accommodate disfluencies like filled pauses and repetitions without fundamentally changing the entire recognition architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If synthetic n-grams with disfluency tokens are added to the language model, then recognition accuracy improves, but memory footprint increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmemory footprint
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by modifying the language model to include synthetic n-grams with disfluency tokens integrated at specific positions within existing n-gram structures. Rather than adding entirely separate data structures, the disfluency tokens are incorporated as parameters within the existing n-gram framework, improving accuracy while controlling memory usage through efficient integration rather than duplication.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10186257B1Language model for speech recognition to account for types of disfluency
Publication Date: 2019.01.22 NVOQ INC
  • US10186257B1 patent drawing
  • US10186257B1 patent drawing
  • US10186257B1 patent drawing

AI summary

The technology of the present application provides a modified language model to allow for the recognition of speech containing types of disfluency. The modified language model includes a plurality of n-grams where an n-gram comprises a sequent of words. The language model also has at least one synthetic n-gram where the synthetic is a naturally occurring n-gram combined with a disfluency token. The disfluency token is representative of multiple types of disfluency and multiple pronunciations thereof.