Automated Transcription Alignment via Phoneme Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated transcription of audio recordings often results in errors due to selecting similar-sounding words, requiring human intervention to correct, and existing methods do not effectively evaluate transcription quality or align automatically generated transcriptions with manually generated ones.

Innovation Solution

A computer-implemented method aligns automatically generated transcriptions with manually generated ones by identifying non-aligned text fragments, mapping phonemes, and using weighted phoneme distances to correct errors, and evaluates transcription quality by computing phoneme distances to detect errors and update models for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated transcription processes are used to transcribe audio recordings, then transcription speed and productivity are improved, but transcription accuracy deteriorates due to errors such as selecting similar-sounding words

Engineering Contradiction:
Improvetranscription speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces phoneme-level analysis as an intermediary between acoustic transcription and final text output. By mapping transcribed words to phonemes and comparing with reference phonemes, the system detects and corrects transcription errors without requiring full manual review, thus maintaining automated speed while improving accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual human proofreading (mechanical human intervention) with an automated phoneme-based error detection and correction system. This substitution maintains high productivity while improving transcription accuracy through computational phoneme comparison and similarity scoring

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual human transcription is used to ensure high transcription accuracy, then transcription quality is improved, but time consumption and productivity deteriorate

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtranscription speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent applies partial action by performing phoneme-level error detection only on transcribed segments that require verification, rather than manually reviewing entire transcriptions. This selective approach maintains high accuracy for critical segments while preserving overall productivity through automated processing of non-critical portions

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If traditional alignment methods are used to align automatically generated transcriptions with manually generated ones, then transcription quality evaluation is achieved, but computational resource consumption increases

Engineering Contradiction:
Improvetranscription quality evaluationVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential phoneme sequences from full transcriptions for comparison and evaluation purposes. By working with phoneme-level representations rather than complete text alignments, the system achieves accurate transcription quality measurement while significantly reducing computational resources required for processing and comparison

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11562743B2Analysis of an automatically generated transcription
Publication Date: 2023.01.24 SALESFORCE INC
  • US11562743B2 patent drawing
  • US11562743B2 patent drawing
  • US11562743B2 patent drawing

AI summary

There is provided a computer implemented method of aligning an automatically generated transcription of an audio recording to a manually generated transcription of the audio recording comprising: identifying non-aligned text fragments, each located between respective two non-continuous aligned text-fragments of the automatically generated transcription, each aligned text-fragment matching words of the manually generated transcription, for each respective non-aligned text fragment: mapping a target keyword of the manually generated transcription to phonemes, mapping the respective non-aligned text fragment to a corresponding audio-fragment of the audio recording, mapping the audio-fragment to phonemes, identifying at least some of the phonemes of the audio-fragment that correspond to the phonemes of the target keyword, and mapping the identified at least some of the phonemes of the audio-fragment to a corresponding word of the automatically generated transcript, wherein the corresponding word is an incorrect automated transcription of the target word appearing in the manually generated transcription.