Conversation Semantic Similarity Measurement Using Edit Distance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for measuring semantic textual similarity are inadequate for comparing conversations due to their unique characteristics, such as multiple authors and ordered flow, which differ from traditional documents.

Innovation Solution

A computer-implemented method that encodes conversation texts into semantic representations and computes minimal edit distance between them, considering costs for deletion, insertion, and substitution operations, with infinite substitution cost for different author types, to quantify semantic similarity and align utterances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional text representation approaches (high dimensional sparse feature vectors) are used for conversations, then computational efficiency is improved, but semantic similarity measurement accuracy deteriorates because traditional methods cannot capture the unique characteristics of conversations (multiple authors, ordered flow, dialog acts)

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidsemantic similarity measurement accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments conversations into individual utterances and represents each utterance separately using contextual embeddings (e.g., BERT). This allows the model to capture the unique characteristics of each utterance while preserving the overall conversation structure, resolving the contradiction between computational efficiency and measurement accuracy by processing conversations in manageable units rather than as monolithic blocks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional high-dimensional sparse feature vectors to dense contextual embeddings that incorporate positional and speaker information dimensions. This dimensional enrichment allows the representation to capture not only semantic meaning but also the structural characteristics of conversations (author identity, turn order), thereby improving measurement accuracy without sacrificing computational efficiency through the use of efficient embedding models.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If contextual representation methods are used for conversations, then semantic similarity measurement accuracy is improved, but device complexity increases due to the need to handle multiple authors, ordered flow, and dialog acts

Engineering Contradiction:
Improvesemantic similarity measurement accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs a universal contextual embedding model (such as BERT) that can simultaneously handle multiple aspects of conversation representation including semantic meaning, speaker identity, and turn order through a single model architecture. This multi-functional approach reduces system complexity compared to using separate models for each aspect, while maintaining high measurement accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges multiple representation elements (semantic embeddings, speaker embeddings, positional embeddings) into a unified conversation representation. By combining these elements through concatenation or summation and then applying a single similarity computation, the system achieves accurate semantic similarity measurement without the complexity of multiple separate processing pipelines.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If traditional document comparison methods are applied to conversations, then ease of operation is maintained, but measurement precision deteriorates because conversations have unique characteristics (multiple authors, ordered flow) that traditional methods cannot capture

Engineering Contradiction:
Improvemethod simplicityVSAvoidconversation similarity measurement accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent creates a structured copy of the conversation data that preserves the unique characteristics (speaker labels, turn order) in a format suitable for computational processing. This copying approach allows traditional comparison algorithms to be applied to the enriched representation without requiring complex modifications to the underlying algorithms, thereby maintaining ease of operation while improving measurement precision through the enriched representation.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11823666B2Automatic measurement of semantic similarity of conversations
Publication Date: 2023.11.21 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11823666B2 patent drawing
  • US11823666B2 patent drawing

AI summary

Automatic measurement of semantic textual similarity of conversations, by: receiving two conversation texts, each comprising a sequence of utterances; encoding each of the sequences of utterances into a corresponding sequence of semantic representations; computing a minimal edit distance between the sequences of semantic representations; and, based on the computation of the minimal edit distance, performing at least one of: quantifying a semantic similarity between the two conversation texts, and outputting an alignment of the two sequences of utterances with each other.