ASR Error Detection via Alternative Result Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition (ASR) systems often produce recognition errors, particularly in low-quality audio or technical contexts, leading to potential serious consequences in domains like medicine due to misrecognition of short sounds or words, which are difficult for human reviewers to detect reliably.

Innovation Solution

A method and apparatus that evaluate ASR results by comparing a top recognition result to alternative results for medically significant differences, using sets of confusable words/phrases and language models to identify potential errors and trigger alerts for review, ensuring accurate transcription in critical domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If ASR systems process speech input to generate recognition results, then speech-to-text conversion is achieved, but recognition errors occur particularly in low-quality audio or technical contexts

Engineering Contradiction:
Improvespeech-to-text conversion efficiencyVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an evaluation engine as an intermediary component between the ASR system and the final output. This evaluation engine receives the recognition result and alternative results from the ASR system, compares them using domain-specific criteria (particularly for medical domains), and determines whether the result contains potential significant errors before passing it to the user interface. This intermediary layer resolves the contradiction by maintaining high conversion efficiency while improving reliability through automated error detection.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If human reviewers manually review ASR results to detect errors, then error detection capability is provided, but the process is time-consuming and difficult to perform reliably

Engineering Contradiction:
Improveerror detection accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements a self-service mechanism where the system automatically evaluates its own recognition results without requiring manual human review. The evaluation engine uses domain-specific knowledge (particularly medical terminology and confusable word pairs) to autonomously identify potential significant errors by comparing the recognition result with alternative results. This self-service approach resolves the contradiction by achieving high error detection accuracy through automated domain-specific analysis while eliminating time loss associated with manual review processes.

Inventive Principle:
Principle #25Self-service

3Reliability

If the system compares top recognition result with alternative results to identify errors, then error detection capability is improved, but system complexity increases

Engineering Contradiction:
Improveerror identification reliabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the error detection function into a separate, dedicated evaluation engine that operates independently from the core ASR recognition process. This segmentation allows the system to maintain simple, efficient speech-to-text conversion while adding a modular error detection component that compares recognition results with alternative results using domain-specific criteria. The segmented architecture resolves the contradiction by improving error identification reliability through specialized comparison logic without significantly increasing overall system complexity, as the evaluation engine operates as an independent module.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9685153B2Detecting potential significant errors in speech recognition results
Publication Date: 2017.06.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9685153B2 patent drawing
  • US9685153B2 patent drawing
  • US9685153B2 patent drawing

AI summary

In some embodiments, the recognition results produced by a speech processing system (which may include a top recognition result and one or more alternative recognition results) based on an analysis of a speech input, are evaluated for indications of potential significant errors. In some embodiments, the recognition results may be evaluated to determine whether a meaning of any of the alternative recognition results differs from a meaning of the top recognition result in a manner that is significant for the domain. In some embodiments, one or more of the recognition results may be evaluated to determine whether the result(s) include one or more words or phrases that, when included in a result, would change a meaning of the result in a manner that would be significant for the domain.