Acoustic Model Training Using Non-Literal Transcripts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional acoustic model training techniques require verbatim transcripts for effective speech recognition, which is challenging in domains like medicine and law due to the difficulty in obtaining large quantities of accurate and domain-specific training data, leading to sub-optimal acoustic models.

Innovation Solution

A system that identifies text representing concepts with multiple spoken forms in non-literal transcripts, replaces this text with context-free grammars to produce a revised transcript, and uses this revised transcript to train acoustic models, improving their accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If verbatim transcripts are used for acoustic model training, then training accuracy is improved, but data availability deteriorates due to the difficulty of obtaining large quantities of accurate domain-specific transcripts

Engineering Contradiction:
Improvetraining accuracyVSAvoiddata availability
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the parameter of transcript fidelity from verbatim to non-verbatim, allowing the use of abundant domain-specific reports that have been modified for readability while still providing sufficient training signal for acoustic model learning

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary process of using non-verbatim transcripts as a bridge between abundant domain reports and the need for accurate training data, enabling training without requiring direct verbatim transcriptions

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If non-literal transcripts are used for training, then data availability is improved, but training effectiveness deteriorates due to deviations from actual speech

Engineering Contradiction:
Improvedata availabilityVSAvoidtraining effectiveness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies partial verbatim transcription, retaining enough speech characteristics in non-verbatim transcripts to provide effective training signal while allowing modifications for readability and domain-specific terminology

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies different levels of verbatim fidelity to different portions of transcripts, maintaining speech accuracy in critical regions while allowing summarization or rephrasing in other areas

Inventive Principle:
Principle #3Local quality

3Measurement precision

If human transcriptionists are used, then transcript accuracy is improved through domain-specific knowledge, but productivity deteriorates due to slow transcription speed

Engineering Contradiction:
Improvetranscript accuracyVSAvoidtranscription speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent enables the acoustic model training system to use domain-specific reports directly with minimal human intervention, allowing the system to self-train on abundant domain data without requiring manual verbatim transcription by experts

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces automated processing as an intermediary between domain reports and training data, using computational methods to bridge the gap between unstructured reports and structured training datasets

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8335688B2Document transcription system training
Publication Date: 2012.12.18 SOLVENTUM INTELLECTUAL PROPERTIES CO
  • US8335688B2 patent drawing
  • US8335688B2 patent drawing
  • US8335688B2 patent drawing

AI summary

A system is provided for training an acoustic model for use in speech recognition. In particular, such a system may be used to perform training based on a spoken audio stream and a non-literal transcript of the spoken audio stream. Such a system may identify text in the non-literal transcript which represents concepts having multiple spoken forms. The system may attempt to identify the actual spoken form in the audio stream which produced the corresponding text in the non-literal transcript, and thereby produce a revised transcript which more accurately represents the spoken audio stream. The revised, and more accurate, transcript may be used to train the acoustic model, thereby producing a better acoustic model than that which would be produced using conventional techniques, which perform training based directly on the original non-literal transcript.