Acoustic Model Training via Transcription Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Acoustic model training for speech recognizers is speaker and channel dependent, requiring manual segmentation of audio recordings which is labor-intensive and prone to inconsistencies, and existing automatic segmentation methods fail to accurately capture sentence or phrase boundaries, leading to inefficient speech training and reduced recognition accuracy.
Innovation Solution
A method that generates conversation-specific language models from transcriptions of multi-channel recordings, aligns written language with audio recordings to determine sentence or phrase boundaries, and trains speech recognizers using these boundaries, eliminating the need for manual segmentation and accounting for speaker and channel variations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual segmentation of audio recordings is used to train speech recognizers, then speaker and channel dependencies can be addressed, but the process becomes labor-intensive and prone to inconsistencies
Solution Approach 1:
The system performs automatic segmentation of audio recordings into sentences and phrases using speech recognition technology and punctuation detection, eliminating the need for manual segmentation. The transcriber's punctuation marks in the transcription automatically define sentence boundaries, which are then used to segment the audio recording without human intervention.
Solution Approach 2:
Manual mechanical segmentation is replaced with an automated computational system that uses speech recognition results, punctuation detection in transcriptions, and audio signal processing to automatically segment recordings into training segments with precise time boundaries.
2Productivity
If existing automatic segmentation methods are used, then productivity is improved, but measurement precision of sentence or phrase boundaries deteriorates
Solution Approach 1:
The transcription with punctuation marks serves as an intermediary that bridges the audio signal and the segmentation process. The punctuation marks in the transcription provide precise linguistic boundaries that guide the segmentation of the audio recording, ensuring both automation and accuracy.
Solution Approach 2:
The transcription is created and punctuated before the segmentation process. These pre-established punctuation marks serve as guides for subsequent automatic segmentation, ensuring that sentence and phrase boundaries are accurately identified without requiring manual intervention during the segmentation phase.
3Measurement precision
If manual segmentation is performed to capture accurate sentence or phrase boundaries, then measurement precision is improved, but device complexity and labor requirements increase
Solution Approach 1:
The system uses its own speech recognition output and the transcriber's punctuation marks to automatically segment the audio, eliminating the need for external manual segmentation tools or expert linguists. The system segments the audio itself based on the punctuation in its own transcription.
Data Source
AI summary
A method, executed by a computer, includes receiving a channel recording corresponding to a conversation, receiving a transcription for the conversation, generating a conversation-specific language model for the conversation using the transcription, and conducting speech recognition on the channel recording using the conversation-specific language model to provide time boundaries and written language corresponding to utterances within the channel recording. The method further includes determining sentence or phrase boundaries for the transcription, aligning written language within the one or more transcriptions with the written language corresponding to the utterances with the channel recording to provide sentence or phrase boundaries for the channel recording, and training a speech recognizer according to the sentence or phrase boundaries for the transcription and the sentence or phrase boundaries for the channel recording. A computer system and computer program product corresponding to the method are also disclosed herein.


