Acoustic Model Training via Transcription Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Acoustic model training for speech recognizers is speaker and channel dependent, requiring manual segmentation of audio recordings which is labor-intensive and prone to inconsistencies, and existing automatic segmentation methods fail to accurately capture sentence or phrase boundaries, leading to inefficient speech training and reduced recognition accuracy.

Innovation Solution

A method that generates conversation-specific language models from transcriptions of multi-channel recordings, aligns written language with audio recordings to determine sentence or phrase boundaries, and trains speech recognizers using these boundaries, eliminating the need for manual segmentation and accounting for speaker and channel variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual segmentation of audio recordings is used to train speech recognizers, then speaker and channel dependencies can be addressed, but the process becomes labor-intensive and prone to inconsistencies

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs automatic segmentation of audio recordings into sentences and phrases using speech recognition technology and punctuation detection, eliminating the need for manual segmentation. The transcriber's punctuation marks in the transcription automatically define sentence boundaries, which are then used to segment the audio recording without human intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical segmentation is replaced with an automated computational system that uses speech recognition results, punctuation detection in transcriptions, and audio signal processing to automatically segment recordings into training segments with precise time boundaries.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing automatic segmentation methods are used, then productivity is improved, but measurement precision of sentence or phrase boundaries deteriorates

Engineering Contradiction:
Improvetraining efficiencyVSAvoidboundary accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The transcription with punctuation marks serves as an intermediary that bridges the audio signal and the segmentation process. The punctuation marks in the transcription provide precise linguistic boundaries that guide the segmentation of the audio recording, ensuring both automation and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The transcription is created and punctuated before the segmentation process. These pre-established punctuation marks serve as guides for subsequent automatic segmentation, ensuring that sentence and phrase boundaries are accurately identified without requiring manual intervention during the segmentation phase.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual segmentation is performed to capture accurate sentence or phrase boundaries, then measurement precision is improved, but device complexity and labor requirements increase

Engineering Contradiction:
Improveboundary accuracyVSAvoidsegmentation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses its own speech recognition output and the transcriber's punctuation marks to automatically segment the audio, eliminating the need for external manual segmentation tools or expert linguists. The system segments the audio itself based on the punctuation in its own transcription.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10096315B2Acoustic model training
Publication Date: 2018.10.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10096315B2 patent drawing
  • US10096315B2 patent drawing
  • US10096315B2 patent drawing

AI summary

A method, executed by a computer, includes receiving a channel recording corresponding to a conversation, receiving a transcription for the conversation, generating a conversation-specific language model for the conversation using the transcription, and conducting speech recognition on the channel recording using the conversation-specific language model to provide time boundaries and written language corresponding to utterances within the channel recording. The method further includes determining sentence or phrase boundaries for the transcription, aligning written language within the one or more transcriptions with the written language corresponding to the utterances with the channel recording to provide sentence or phrase boundaries for the channel recording, and training a speech recognizer according to the sentence or phrase boundaries for the transcription and the sentence or phrase boundaries for the channel recording. A computer system and computer program product corresponding to the method are also disclosed herein.