Hybrid Voice Messaging System with Human-in-the-Loop Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice messaging systems struggle to efficiently convert unstructured voicemail messages into text at a mass scale, particularly for large user bases, due to high costs and inefficiencies in processing times, and fail to accurately capture the meaning and idiomatic elements of messages, which is crucial for user confidence and practical application.

Innovation Solution

A hybrid system combining automatic speech recognition (ASR) with human operators and quality control, utilizing pre-processing, conversion resources, and quality control sub-systems to optimize human operator efficiency, and employing contextual information and language models to improve conversion accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional ASR systems are used to convert voicemail to text, then transcription can be automated, but the system fails to capture meaning and idiomatic elements accurately

Engineering Contradiction:
Improveautomated transcriptionVSAvoidaccuracy of meaning capture
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces human operators as an intermediary between the automated ASR system and the final text output. The ASR system provides initial transcription, which is then reviewed and refined by human operators who understand context, idioms, and meaning, thus combining automation with semantic accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a composite transcription process combining machine-generated text from ASR with human editorial input. This composite approach leverages the speed and scalability of automation while incorporating human understanding of nuance, resulting in text that captures both efficiency and meaning

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If human operators transcribe all messages, then accuracy improves, but cost and processing time increase prohibitively for mass scale

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing speed at scale
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system applies human review selectively rather than universally. Human operators review and refine only the portions of transcriptions that require contextual understanding or contain uncertain ASR output, while accepting straightforward transcriptions as-is, thus optimizing the balance between accuracy and processing capacity

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The transcription process is segmented into multiple stages: initial automated ASR transcription, quality assessment of the ASR output, selective human review of problematic segments, and final assembly. This segmentation allows the system to handle mass volume while maintaining accuracy where needed

Inventive Principle:
Principle #1Segmentation

3Loss of time

If the system processes messages quickly to meet 2-5 minute turnaround, then user confidence improves, but processing complexity increases for large user bases

Engineering Contradiction:
Improvemessage turnaround timeVSAvoidsystem processing complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system performs preliminary automated ASR transcription immediately upon receiving a voicemail, creating a draft text version before human review. This preliminary action ensures that even if human review takes time, a usable transcription is already available, meeting the fast turnaround requirement while preparing the message for subsequent quality enhancement

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous processing through parallel operations: ASR transcription occurs continuously for all incoming messages, while human operators continuously review and refine transcriptions. This continuous multi-stage processing ensures consistent fast turnaround times even as message volume scales to large user bases

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9191515B2Mass-scale, user-independent, device-independent voice messaging system
Publication Date: 2015.11.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9191515B2 patent drawing
  • US9191515B2 patent drawing
  • US9191515B2 patent drawing

AI summary

A mass-scale, user-independent, device-independent, voice messaging system that converts unstructured voice messages into text for display on a screen is disclosed. The system comprises (i) computer implemented sub-systems and also (ii) a network connection to human operators providing transcription and quality control; the system being adapted to optimize the effectiveness of the human operators by further comprising 3 core sub-systems, namely (i) a pre-processing front end that determines an appropriate conversion strategy; (ii) one or more conversion resources; and (iii) a quality control sub-system.