Agent-Assisted Transcription with Context-Aware Accelerants

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems often fail to accurately transcribe user utterances in real-time, necessitating human agent intervention which can be slow and inefficient, leading to delays in online system responses.

Innovation Solution

Implementing an agent-assisted transcription method where a language model is dynamically selected based on context and user information, with human agents reviewing and correcting automatic transcriptions, and using text predictions to enhance transcription speed and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automatic speech recognition is used for transcription, then transcription speed is improved, but transcription accuracy deteriorates

Engineering Contradiction:
Improvetranscription speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces human agents as intermediaries between automatic speech recognition and the final transcription output. Agents review ASR transcriptions and correct errors, combining the speed of automated systems with the accuracy of human judgment. The system presents ASR results to agents who then refine the transcription quality through manual correction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent merges automatic speech recognition technology with human agent review into a hybrid transcription system. This combination leverages the strengths of both approaches: ASR provides rapid initial transcription while human agents provide accuracy verification and correction, achieving both speed and precision.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If human agents are used for transcription correction, then transcription accuracy is improved, but response time deteriorates

Engineering Contradiction:
Improvetranscription accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automatic speech recognition transcription before human agent review. This preliminary action creates a draft transcription that agents can quickly review and correct, rather than starting from scratch. The pre-processing by ASR reduces the workload on agents and accelerates the overall process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses ASR to create a copy or draft of the transcription that agents then review and correct. This copying approach allows agents to work with pre-generated content rather than creating transcriptions manually, significantly reducing response time while maintaining accuracy through selective human review.

Inventive Principle:
Principle #26Copying

3Measurement precision

If context-specific language models are used, then transcription accuracy is improved, but system complexity deteriorates

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system dynamically changes language model parameters based on context. Different language models are selected depending on the specific transcription context, such as domain-specific terminology or speaker characteristics. This parameter adjustment improves accuracy without requiring a completely complex system architecture, as it leverages existing models with different configurations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11314942B1Accelerating agent performance in a natural language processing system
Publication Date: 2022.04.26 INTERACTIONS LLC (US)
  • US11314942B1 patent drawing
  • US11314942B1 patent drawing
  • US11314942B1 patent drawing

AI summary

A computer-implemented method for providing agent assisted transcriptions of user utterances. A user utterance is received in response to a prompt provided to the user at a remote client device. An automatic transcription is generated from the utterance using a language model based upon an application or context, and presented to a human agent. The agent reviews the transcription and may replace at least a portion of the transcription with a corrected transcription. As the agent inputs the corrected transcription, accelerants are presented to the user comprising suggested texted to be inputted. The accelerants may be determined based upon an agent input, an application or context of the transcription, the portion of the transcription being replaced, or any combination thereof. In some cases, the user provides textual input, to which the agent transcribes an intent associated with the input with the aid of one or more accelerants.