Conversational Speech Summarization via Segmentation and Quality Profiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in effectively summarizing conversational speech due to variations in speech patterns, mid-conversation context switches, domain-specific vocabulary, and culturally specific terms, making it costly or impossible to train machine learning models for gaining insights from captured conversations.

Innovation Solution

The development of a system that segments conversational speech using segmentation models, generates summarization profiles based on these segments, and updates models with quality scores to provide accurate summaries, including metadata and metrics such as topic adherence and speaker information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are trained to provide insights from captured conversations, then the quality of speech analysis and summarization improves, but the cost and complexity of training these models becomes prohibitively expensive or impossible

Engineering Contradiction:
Improvespeech analysis qualityVSAvoidmodel training cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments conversational speech into distinct segments based on context switches, topic changes, and speaker transitions. This segmentation allows the system to process and analyze speech in manageable units rather than attempting to train complex models on entire conversations, thereby improving analysis quality while reducing the computational cost and complexity associated with model training.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If comprehensive machine transcription services are used to transform natural speech into text, then the completeness of speech capture improves, but the resources required for transcription and analysis increase significantly

Engineering Contradiction:
Improvespeech capture completenessVSAvoidtranscription efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts only the essential and relevant portions of conversational speech by segmenting conversations based on context switches and topic changes. This extraction approach allows the system to capture the most important information without processing every single word of the entire conversation, thereby maintaining speech capture completeness while significantly improving transcription efficiency and reducing resource requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If machine learning models are trained to handle variations in speech patterns, context switches, and domain-specific vocabulary, then the accuracy of speech analysis improves, but the training becomes prohibitively expensive or impossible

Engineering Contradiction:
Improvespeech analysis accuracyVSAvoidmodel training feasibility
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent divides conversational speech into distinct segments based on context switches, topic changes, and speaker transitions. This segmentation enables the system to handle variations in speech patterns and domain-specific vocabulary by processing each segment independently with appropriate analysis methods, thereby maintaining high accuracy without requiring prohibitively expensive comprehensive model training.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different analysis approaches to different segments of conversation based on their specific characteristics. By recognizing context switches and topic changes, the system can apply domain-specific vocabulary handling and speech pattern analysis tailored to each segment, improving overall accuracy without needing a single comprehensive model trained on all possible variations.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11842144B1Summarizing conversational speech
Publication Date: 2023.12.12 INVOCA INC
  • US11842144B1 patent drawing
  • US11842144B1 patent drawing
  • US11842144B1 patent drawing

AI summary

Embodiments are directed to summarizing conversational speech. Conversation segments may be provided based on a conversation stream and segmentation models. Summarization models may be determined based on characteristics of the conversation segments. Summarization information may be generated for each of the conversation segments based on the summarization models such that the summarization information includes a text-based summarization of the conversation segment. Summarization profiles may be generated for the conversation segments based on the summarization information such that each summarization profile is associated with quality scores. Summarization models may be modified based on the summarization profiles and the associated quality scores such that the summarization profiles are updated based on the modified summarization models. Modified summarization models and the updated summarization profiles may be employed to provide reports to a user.