Adaptive Text Prediction Using Dynamic Source Configuration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Human-based speech-to-text transcription is costly and often of poor quality due to time constraints and variable audio quality, while existing machine-based solutions lack consideration for linguistic rules and context, resulting in unsatisfactory transcription results.

Innovation Solution

A computerized method that determines the configuration of multiple prediction sources, including language models and human agents, based on features of the voice data to generate an adaptive textual prediction, optimizing the order and weighting of these sources for improved transcription quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human agents are used for speech-to-text transcription, then transcription quality can be maintained, but cost increases and time constraints reduce efficiency

Engineering Contradiction:
Improvetranscription qualityVSAvoidtranscription efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts the configuration of prediction sources based on audio features such as signal-to-noise ratio and lattice complexity. When audio quality is high and lattice complexity is low, the system relies more on automated prediction sources. When audio quality degrades or complexity increases, the system adaptively incorporates human agent intervention, thereby optimizing both efficiency and quality according to real-time conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters by adjusting the weighting and configuration of different prediction sources (automated language models vs. human agents) based on extracted audio features. This parameter adjustment allows the system to maintain high transcription quality while reducing reliance on human agents in favorable conditions, thus improving overall productivity.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If machine-based speech recognition is used, then cost and efficiency are improved, but transcription quality deteriorates due to lack of linguistic context understanding

Engineering Contradiction:
Improvetranscription efficiencyVSAvoidtranscription quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system merges multiple prediction sources including automated language models and human agent capabilities into a unified transcription system. By combining the efficiency of machine-based recognition with the contextual understanding and linguistic expertise of human agents, the system achieves both high productivity and high transcription quality simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an adaptive configuration layer that acts as an intermediary between automated speech recognition and final transcription output. This intermediary dynamically selects and weights different prediction sources based on audio features, allowing machine-based efficiency to be maintained while compensating for quality deficiencies through selective human agent involvement.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multiple prediction sources are used, then transcription accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem configuration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the transcription task by dividing it into multiple prediction sources with different specializations (e.g., language models for general text, human agents for complex linguistic contexts). Each prediction source handles specific aspects of the transcription task, improving overall accuracy while the modular segmented structure helps manage system complexity through clear division of responsibilities.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9099091B2Method and apparatus of adaptive textual prediction of voice data
Publication Date: 2015.08.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9099091B2 patent drawing
  • US9099091B2 patent drawing
  • US9099091B2 patent drawing

AI summary

Typical textual prediction of voice data employs a predefined implementation arrangement of a single or multiple prediction sources. Using a predefined implementation arrangement of the prediction sources may not provide a good prediction performance in a consistent manner with variations in voice data quality. Prediction performance may be improved by employing adaptive textual prediction. According to at least one embodiment determining a configuration of a plurality of prediction sources, used for textual interpretation of the voice data, is determined based at least in part on one or more features associated with the voice data or one or more a-priori interpretations of the voice data. A textual output prediction of the voice data is then generated using the plurality of prediction sources according to the determined configuration. Employing an adaptive configuration of the text prediction sources facilitates providing more accurate text transcripts of the voice data.