Chat Transcript Analysis via Utterance Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current chat transcript analysis tools fail to provide a complete understanding of online chat data, requiring manual analysis of large datasets and often resulting in anecdotal conclusions based on small subsets of data.

Innovation Solution

A method that classifies utterances from chat transcripts into products and intents, applies tags to categorize the data, and clusters them based on sentence similarity to extract representative utterances with the highest semantic similarity, providing a more accurate and efficient understanding of the underlying meaning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If chat transcript analysis tools are used to summarize conversation topics, then analysis speed is improved, but understanding completeness deteriorates

Engineering Contradiction:
Improveanalysis speedVSAvoidunderstanding completeness
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments chat transcript analysis into multiple classification dimensions: product classification, intent classification, and cluster classification. Each dimension provides specific insights (product discussions, customer intentions, topic groupings) that collectively deliver complete understanding while maintaining automated processing speed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds multiple classification dimensions (product, intent, cluster) to the traditional single-topic summarization approach. This multi-dimensional classification framework enables comprehensive understanding of chat transcripts by capturing different aspects of customer interactions simultaneously, resolving the trade-off between speed and completeness.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If manual analysis of raw transcript data is performed, then understanding completeness is improved, but analysis time increases

Engineering Contradiction:
Improveunderstanding completenessVSAvoidanalysis time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements automated self-service classification where the system automatically classifies utterances into products, intents, and clusters without human intervention. The multi-dimensional classification framework processes large volumes of chat transcript data autonomously, delivering comprehensive insights at automated processing speeds, thus eliminating the time cost of manual analysis while maintaining understanding completeness.

Inventive Principle:
Principle #25Self-service

3Loss of time

If manual analysis is performed on small subsets of data, then analysis time is reduced, but result reliability deteriorates

Engineering Contradiction:
Improveanalysis timeVSAvoidresult reliability
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent changes the parameter of data processing scale by enabling automated multi-dimensional classification of entire chat transcript datasets. This allows comprehensive analysis of large datasets at automated processing speeds, producing statistically reliable results that generalize well to the overall customer base, unlike anecdotal conclusions from small manual samples.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11170173B2Analyzing chat transcript data by classifying utterances into products, intents and clusters
Publication Date: 2021.11.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11170173B2 patent drawing
  • US11170173B2 patent drawing
  • US11170173B2 patent drawing

AI summary

A method, system and computer program product for improving the understanding of chat transcript data. Chat transcripts are analyzed to classify the utterances into intents and identify products discussed in the chat transcripts. The data of the chat transcripts are divided into categories of utterances associated with products and intents by applying tags to the chat transcripts. The categories of utterances associated with products and intents are then clustered into clusters based on sentence similarity. Once the utterances are grouped, a representative utterance is extracted from a cluster, where the representative utterance is an utterance that has the highest semantic similarity to the utterances in the cluster. In this manner, users will be provided a more accurate guide as to the underlying meaning of the chat transcript data thereby improving the understanding of the chat transcript data more efficiently and accurately than current chat transcript analysis tools.