Confusion Index for Speech Recognition Word Conflict Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current acoustic-based approaches for measuring word confusion in language models and ASR systems treat all similar-sounding words equally, failing to differentiate levels of confusion and their impact on model performance, leading to inaccurate recognition.
Innovation Solution
A system and method that calculate a confusion index by combining acoustic and language understanding, using acoustic distance and language scores to classify and handle different types of word confusions, with a weighting factor to balance their influence, allowing for context addition or boosting of words to resolve conflicts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If acoustic-based approaches are used to measure word confusion, then similar-sounding words can be identified, but all such words are treated equally without differentiating levels of confusion
Solution Approach 1:
The patent segments the confusion measurement into multiple dimensions: acoustic similarity (phonetic distance) and language model probability. By dividing the measurement into these separate components, the system can identify similar-sounding words while also differentiating their confusion levels based on contextual probability, thus resolving the contradiction between identification capability and measurement complexity.
Solution Approach 2:
The patent introduces a confusion index that combines acoustic distance parameters with language model probability parameters. This parameter change transforms the measurement from a single acoustic metric to a composite index that differentiates between various levels of confusion, allowing the system to treat similar-sounding words differently based on their contextual likelihood.
2Ease of manufacture
If acoustic similarity alone is used to measure confusion, then the measurement process is simple, but it does not provide insight into the degree or level of confusion and its impact on models
Solution Approach 1:
The patent creates a composite measurement approach by combining acoustic similarity data with language model probability data. This composite confusion index retains the simplicity of acoustic measurement while incorporating additional linguistic information to provide insight into the degree and impact of confusion, thus preventing information loss while maintaining measurement feasibility.
3Ease of operation
If all similar-sounding words are treated alike, then the handling process is straightforward, but it does not enable intelligent handling to reduce negative effects on language models
Solution Approach 1:
The patent applies local quality by treating different similar-sounding words differently based on their specific confusion characteristics. Instead of uniform handling, the system calculates individual confusion indices for each word pair and applies targeted resolution strategies, such as boosting language model probabilities for specific words, thereby improving recognition accuracy while maintaining operational efficiency.
Data Source
AI summary
Systems and methods to improve the performance of an automatic speech recognition (ASR) system using a confusion index indicative of the amount of confusion between words are described, where a confusion index (CI) or score is calculated by receiving a first word (Word1) and a second word (Word2), calculating an acoustic score (A12) indicative of the phonetic difference between Word1 and Word2, calculating a weighted language score (W(U1+U2), indicative of a weighted likelihood (or word frequency) of Word1 and Word2 occurring in the corpus, the confusion index CI incorporating both the acoustic score and the weighted language score, such that the CI for words that sound alike and have a high likelihood of occurring in the corpus will be higher than the CI for words that sound alike and do not have a high likelihood of occurring in the corpus. In some embodiments, the CI may be used to artificially boost uncommon words in a corpus to improve their visibility, to add context to uncommon words in a corpus to avoid conflict with common words, and to remove unimportant words from the lexicon to avoid conflicts with other corpus words.


