Chat Utterance Disentanglement via Linguistic Drift Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing systems fail to effectively disentangle multiple discourses related to different topics during real-time online chat conversations, as they cannot identify and separate parallel discourses or provide users with a way to disentangle entangled chat utterances.

Innovation Solution

A computer-implemented method using corpus linguistics and topic modeling analyzes linguistic collocations and keywords to determine the level of drift in chat utterances, employing an utterance location adjustment score to disentangle chat utterances related to different topics by removing statistically significant drift and rearranging them into new discourses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If corpus linguistics and topic modeling are used to analyze chat utterances, then the ability to identify and separate parallel discourses is improved, but the computational complexity and processing time increase

Engineering Contradiction:
Improvediscourse identification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the chat conversation into multiple parallel discourses by analyzing linguistic collocations and keywords. Each discourse is identified as a separate topic thread, allowing the system to handle complex multi-topic conversations by breaking them down into manageable segments that can be processed independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary computational layer that uses corpus linguistics techniques (collocation analysis, keyword extraction) and topic modeling to bridge the raw chat data and the final discourse identification. This intermediary processing layer transforms unstructured chat utterances into structured discourse representations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If multiple parallel discourses are disentangled and separated, then user understanding and clarity are improved, but the system complexity and processing requirements increase

Engineering Contradiction:
Improveinformation clarityVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent adds a new dimension to chat analysis by introducing discourse-level organization alongside traditional message sequencing. Instead of only temporal ordering, the system creates a discourse structure dimension that groups messages by topic, allowing users to navigate conversations through multiple organizational perspectives simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent performs preliminary analysis of linguistic patterns, collocations, and keywords before full discourse identification. By pre-processing the chat data to extract meaningful linguistic features and establish baseline topic representations, the system reduces the computational burden of subsequent discourse separation and organization.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If real-time analysis of linguistic collocations and keywords is performed, then discourse separation accuracy is improved, but the processing speed decreases

Engineering Contradiction:
Improvediscourse separation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent applies partial analysis by focusing on key linguistic features (collocations and keywords) rather than analyzing every aspect of each message. This selective approach extracts the most informative elements for discourse identification while reducing overall processing requirements and improving speed.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent dynamically adjusts analysis parameters such as collocation window size, keyword threshold, and discourse similarity thresholds based on conversation characteristics. By adapting these parameters in real-time, the system optimizes the balance between analysis depth and processing speed for different types of chat interactions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11449683B2Disentanglement of chat utterances
Publication Date: 2022.09.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11449683B2 patent drawing
  • US11449683B2 patent drawing
  • US11449683B2 patent drawing

AI summary

Disentanglement of chat utterances is provided. An analysis of the linguistic collocations and the keywords of the multiple chat utterances and amount of contribution by respective chat users of the plurality of chat users to the multiple chat utterances is performed to determine a level of drift of the linguistic collocations, the keywords, and respective chat users over a course of the multiple chat utterances. Chat utterance entanglement of prior chat utterances is determined using determined level of drift based on the analysis by inferring keyword usage over time and how these keywords are related over the course of the multiple chat utterances. The prior chat utterances related to a particular topic are disentangled by removing certain chat utterances that have a statistically significant level of drift from that particular topic. Removed chat utterances are arranged as a new chat discourse related to a different topic in the chat conversation.