Speech Recognition Corpus Error Detection Using Perplexity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting transcription errors in speech recognition corpora are inadequate, as they either focus on speech recognition results alone or ignore text characteristics, leading to unreliable learning outcomes for speech recognition models.

Innovation Solution

A method and device that utilize a speech recognition model and a language model to extract performance evaluation indices and perplexities, automatically detecting transcription errors by setting primary and secondary error candidates based on predefined reference values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing error detection methods are used, then the process is simple, but the detection accuracy is insufficient because they only consider speech recognition results without considering text characteristics

Engineering Contradiction:
Improvetranscription error detection accuracyVSAvoiderror detection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines speech recognition results with language model predictions to create a comprehensive error detection system. By merging the speech recognition output with language model probability assessments, the system achieves higher detection accuracy while maintaining reasonable complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The language model acts as an intermediary between the speech recognition system and the error detection process. It provides probability scores that mediate the comparison between transcribed text and expected language patterns, enabling accurate error detection without directly analyzing raw speech signals.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual inspection of large corpora is performed, then detection accuracy is high, but time consumption increases significantly

Engineering Contradiction:
Improvecorpus inspection efficiencyVSAvoidtime for manual inspection
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically detecting transcription errors using the integrated speech recognition and language model approach. This eliminates the need for manual inspection while maintaining high detection accuracy, thereby significantly reducing time consumption and improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical inspection with an automated computational system. By substituting human reviewers with an automated pipeline that combines speech recognition and language model analysis, the system achieves scalable error detection without time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If transcription errors are not detected, then the corpus can be used as-is, but the reliability of speech recognition model learning is compromised

Engineering Contradiction:
Improvelearning result reliabilityVSAvoidcorpus preparation effort
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system performs preliminary action by automatically detecting and identifying transcription errors before the corpus is used for training. This preliminary error detection ensures that only reliable data is used for learning, thereby improving reliability without requiring extensive manual corpus preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The language model provides feedback on the quality of transcriptions by comparing predicted probabilities against actual transcribed text. This feedback mechanism enables automatic identification of errors, ensuring reliable learning data while simplifying corpus preparation efforts.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12431137B2Method of detecting a transcription error in speech recognition corpus and device for the same
Publication Date: 2025.09.30 SOGANG UNIV RES & BUSINESS DEV FOUND
  • US12431137B2 patent drawing
  • US12431137B2 patent drawing

AI summary

Provided are a method and device for detecting transcription error in a speech recognition corpus. The method of detecting a transcription error in speech recognition corpus includes following steps: (a) receiving the speech recognition corpus including a speech file and a text label for the speech file; (b) performing speech recognition on the speech file of the speech recognition corpus using a speech recognition model and converting the speech recognition result into text; (c) extracting a performance evaluation index of the speech recognition model; (d) extracting a PPL(s2) for the text label and a PPL(s1) for the text using a language model; and (e) detecting a transcription error in text label of the speech recognition corpus using the extracted performance evaluation index and the PPL(s2) and PPL(s1).