Automated Transcription Alignment via Phoneme Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated transcription of audio recordings often results in errors due to selecting similar-sounding words, requiring human intervention to correct, and existing methods do not effectively evaluate transcription quality or align automatically generated transcriptions with manually generated ones.
Innovation Solution
A computer-implemented method aligns automatically generated transcriptions with manually generated ones by identifying non-aligned text fragments, mapping phonemes, and using weighted phoneme distances to correct errors, and evaluates transcription quality by computing phoneme distances to detect errors and update models for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated transcription processes are used to transcribe audio recordings, then transcription speed and productivity are improved, but transcription accuracy deteriorates due to errors such as selecting similar-sounding words
Solution Approach 1:
The patent introduces phoneme-level analysis as an intermediary between acoustic transcription and final text output. By mapping transcribed words to phonemes and comparing with reference phonemes, the system detects and corrects transcription errors without requiring full manual review, thus maintaining automated speed while improving accuracy
Solution Approach 2:
The patent replaces manual human proofreading (mechanical human intervention) with an automated phoneme-based error detection and correction system. This substitution maintains high productivity while improving transcription accuracy through computational phoneme comparison and similarity scoring
2Manufacturing precision
If manual human transcription is used to ensure high transcription accuracy, then transcription quality is improved, but time consumption and productivity deteriorate
Solution Approach 1:
The patent applies partial action by performing phoneme-level error detection only on transcribed segments that require verification, rather than manually reviewing entire transcriptions. This selective approach maintains high accuracy for critical segments while preserving overall productivity through automated processing of non-critical portions
3Measurement precision
If traditional alignment methods are used to align automatically generated transcriptions with manually generated ones, then transcription quality evaluation is achieved, but computational resource consumption increases
Solution Approach 1:
The patent extracts only the essential phoneme sequences from full transcriptions for comparison and evaluation purposes. By working with phoneme-level representations rather than complete text alignments, the system achieves accurate transcription quality measurement while significantly reducing computational resources required for processing and comparison
Data Source
AI summary
There is provided a computer implemented method of aligning an automatically generated transcription of an audio recording to a manually generated transcription of the audio recording comprising: identifying non-aligned text fragments, each located between respective two non-continuous aligned text-fragments of the automatically generated transcription, each aligned text-fragment matching words of the manually generated transcription, for each respective non-aligned text fragment: mapping a target keyword of the manually generated transcription to phonemes, mapping the respective non-aligned text fragment to a corresponding audio-fragment of the audio recording, mapping the audio-fragment to phonemes, identifying at least some of the phonemes of the audio-fragment that correspond to the phonemes of the target keyword, and mapping the identified at least some of the phonemes of the audio-fragment to a corresponding word of the automatically generated transcript, wherein the corresponding word is an incorrect automated transcription of the target word appearing in the manually generated transcription.


