Real-Time Speech Transcription Using Context-Based Model Assignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current customer service call centers face inefficiencies and loss of contextual information due to difficulties in consumer interactions with automated and semi-automated systems, leading to high costs and ineffective data harvesting.

Innovation Solution

A computer-implemented method and system that electronically monitors telephonic interactions between consumers and service providers, assigning context-based speech recognition models in real-time to transcribe speech into text, using hierarchical models that adapt based on language, dialect, and transaction roles, and swapping models as necessary to improve accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated and semi-automated systems are used for customer service interactions, then operational efficiency is improved, but contextual information is lost and interaction accuracy deteriorates

Engineering Contradiction:
Improveoperational efficiencyVSAvoidcontextual information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent replaces traditional automated voice recognition systems with an AI-powered natural language processing system that can understand contextual information, language nuances, and dialect variations, thereby maintaining operational efficiency while preventing loss of contextual information

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system dynamically adjusts speech recognition parameters based on detected language, dialect, and transaction role, allowing it to adapt to different interaction contexts and maintain high accuracy across diverse customer service scenarios

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If context-based speech recognition models are assigned to different channels, then transcription accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition system into multiple specialized models organized in a hierarchy, with each model trained for specific languages, dialects, or transaction roles, allowing high accuracy for each segment while managing overall system complexity through modular organization

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects and switches between different speech recognition models based on the detected context, language, and transaction role, enabling adaptive accuracy improvement without requiring all models to be active simultaneously, thus managing computational resources efficiently

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If hierarchical speech recognition models are used with multiple variants, then adaptability to different languages and roles is improved, but processing time increases

Engineering Contradiction:
Improvelanguage and role adaptabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent pre-organizes speech recognition models in a hierarchical structure with language, dialect, and role variants ready for selection, and performs rapid context analysis at the beginning of each interaction to determine the appropriate model, avoiding time-consuming model switching during the interaction

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically detects the language, dialect, and transaction role during the interaction and self-selects the appropriate speech recognition model without requiring manual intervention or extensive processing, enabling rapid adaptation to different contexts

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11114092B2Real-time voice processing systems and methods
Publication Date: 2021.09.07 GRP ALLO MEDIA SAS
  • US11114092B2 patent drawing
  • US11114092B2 patent drawing
  • US11114092B2 patent drawing

AI summary

A computer-implemented method and supporting system transcribes spoken words being monitored from a telephonic interaction among two or more individuals. Telephonic interactions among the individuals are monitored, and at least two of the individuals are each assigned to a separate channel. While still being monitored, each of the channels is assigned a context-based speech recognition models, and in substantially real-time, the monitored telephonic interaction is transcribed from speech to text based on the different assigned models.