Modified Language Model for Disfluency Handling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face limitations in handling disfluencies, such as filled pauses, repetitions, and repairs, which lead to inefficiencies and errors in free form, continuous speech recognition due to their inability to account for these speech patterns in language models.
Innovation Solution
A modified language model is introduced that includes synthetic n-grams, which incorporate disfluency tokens to improve the recognition of speech by adapting the existing clean language model to account for disfluencies, allowing for more accurate matching of phoneme sequences and reducing errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a clean language model is used for speech recognition, then the system is simpler and faster to deploy, but it cannot handle disfluencies in free form speech
Solution Approach 1:
The patent applies preliminary action by pre-generating synthetic n-grams that incorporate disfluency tokens and integrating them into the language model before deployment. This allows the system to handle disfluencies without requiring complex real-time processing during speech recognition, as the disfluency patterns are already accounted for in the pre-built language model structure.
Solution Approach 2:
The patent introduces disfluency tokens as intermediary elements that bridge clean speech patterns and actual disfluent speech. These tokens act as mediators within the language model, allowing the system to recognize and accommodate disfluencies like filled pauses and repetitions without fundamentally changing the entire recognition architecture.
2Measurement precision
If synthetic n-grams with disfluency tokens are added to the language model, then recognition accuracy improves, but memory footprint increases
Solution Approach 1:
The patent applies parameter changes by modifying the language model to include synthetic n-grams with disfluency tokens integrated at specific positions within existing n-gram structures. Rather than adding entirely separate data structures, the disfluency tokens are incorporated as parameters within the existing n-gram framework, improving accuracy while controlling memory usage through efficient integration rather than duplication.
Data Source
AI summary
The technology of the present application provides a modified language model to allow for the recognition of speech containing types of disfluency. The modified language model includes a plurality of n-grams where an n-gram comprises a sequent of words. The language model also has at least one synthetic n-gram where the synthetic is a naturally occurring n-gram combined with a disfluency token. The disfluency token is representative of multiple types of disfluency and multiple pronunciations thereof.


