Machine-Learning Transcript Tagging and Real-Time Whispers for Call Insights

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer-based text/speech analyzers fail to properly capture insights, such as conclusions and analyses, from interactive communications.

Innovation Solution

A system and method that utilizes machine learning models to analyze audio and textual versions of interactive communications, tagging insights in real-time and overlaying them as audio or textual descriptions for immediate assistance to call agents or managers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are used to analyze audio and text in real-time, then insight detection capability is improved, but system complexity increases

Engineering Contradiction:
Improveinsight detection capabilityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex analysis task into separate components: audio processing, text processing, and insight generation. Each component handles a specific aspect of the communication, allowing the system to manage complexity through modular architecture while maintaining high detection precision through specialized processing for each modality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that bridges raw audio/text input and insight output. This intermediary layer includes components for transcription, sentiment analysis, and topic detection that transform complex unstructured data into structured insights, reducing overall system complexity while improving detection capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If real-time analysis is implemented, then response speed is improved, but computational resource consumption increases

Engineering Contradiction:
Improveresponse speedVSAvoidcomputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system implements periodic processing where audio and text are analyzed at specific intervals rather than continuously. This allows real-time response capability while reducing computational load by processing data in discrete batches, balancing speed requirements with resource consumption constraints.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent applies partial processing by focusing computational resources on detecting specific types of insights (sentiment, topics, key phrases) rather than analyzing all aspects of the communication equally. This selective approach maintains response speed for critical insights while reducing overall computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12374324B2Transcript tagging and real-time whisper in interactive communications
Publication Date: 2025.07.29 CAPITAL ONE SERVICES LLC
  • US12374324B2 patent drawing
  • US12374324B2 patent drawing
  • US12374324B2 patent drawing

AI summary

Disclosed herein are system, method, and computer readable medium embodiments for machine learning systems to process interactive communications between at least two participants. Speech and text within the interactive communications are analyzed using machine learning models to infer insights located within the interactive communications. The inferred insights are converted to descriptive text or audio and tagged to the interactive communication as graphics or audio whispers reflecting the insights added to the interactive communication.