Agent-Assisted Transcription with Context-Aware Accelerants
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition systems often fail to accurately transcribe user utterances in real-time, necessitating human agent intervention which can be slow and inefficient, leading to delays in online system responses.
Innovation Solution
Implementing an agent-assisted transcription method where a language model is dynamically selected based on context and user information, with human agents reviewing and correcting automatic transcriptions, and using text predictions to enhance transcription speed and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automatic speech recognition is used for transcription, then transcription speed is improved, but transcription accuracy deteriorates
Solution Approach 1:
The patent introduces human agents as intermediaries between automatic speech recognition and the final transcription output. Agents review ASR transcriptions and correct errors, combining the speed of automated systems with the accuracy of human judgment. The system presents ASR results to agents who then refine the transcription quality through manual correction.
Solution Approach 2:
The patent merges automatic speech recognition technology with human agent review into a hybrid transcription system. This combination leverages the strengths of both approaches: ASR provides rapid initial transcription while human agents provide accuracy verification and correction, achieving both speed and precision.
2Measurement precision
If human agents are used for transcription correction, then transcription accuracy is improved, but response time deteriorates
Solution Approach 1:
The system performs preliminary automatic speech recognition transcription before human agent review. This preliminary action creates a draft transcription that agents can quickly review and correct, rather than starting from scratch. The pre-processing by ASR reduces the workload on agents and accelerates the overall process.
Solution Approach 2:
The system uses ASR to create a copy or draft of the transcription that agents then review and correct. This copying approach allows agents to work with pre-generated content rather than creating transcriptions manually, significantly reducing response time while maintaining accuracy through selective human review.
3Measurement precision
If context-specific language models are used, then transcription accuracy is improved, but system complexity deteriorates
Solution Approach 1:
The system dynamically changes language model parameters based on context. Different language models are selected depending on the specific transcription context, such as domain-specific terminology or speaker characteristics. This parameter adjustment improves accuracy without requiring a completely complex system architecture, as it leverages existing models with different configurations.
Data Source
AI summary
A computer-implemented method for providing agent assisted transcriptions of user utterances. A user utterance is received in response to a prompt provided to the user at a remote client device. An automatic transcription is generated from the utterance using a language model based upon an application or context, and presented to a human agent. The agent reviews the transcription and may replace at least a portion of the transcription with a corrected transcription. As the agent inputs the corrected transcription, accelerants are presented to the user comprising suggested texted to be inputted. The accelerants may be determined based upon an agent input, an application or context of the transcription, the portion of the transcription being replaced, or any combination thereof. In some cases, the user provides textual input, to which the agent transcribes an intent associated with the input with the aid of one or more accelerants.


