Trending Topic Identification via Cluster Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for identifying trending topics from textual data streams often generate multiple unwanted alerts for the same topic due to variations in terminology over time, leading to inefficiencies in alert management.

Innovation Solution

A system that receives text documents, identifies abnormal terms by frequency analysis, clusters terms based on co-occurrence, and compares clusters across time periods to determine if they represent the same topic, thereby filtering out redundant alerts by associating each term with document identifiers and time information and using abnormality scores to update models and reduce duplicate alerts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If standard frequency-based abnormal term identification is used, then new trending topics can be detected, but multiple unwanted alerts are generated for the same topic due to terminology variations

Engineering Contradiction:
Improveaccuracy of topic identificationVSAvoidspurious alerts
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent combines multiple clusters that represent the same topic by comparing clusters across different time periods. When a new cluster is formed, it is compared with existing clusters to determine if they refer to the same topic, and if so, they are merged into a single alert. This resolves the contradiction by maintaining reliable topic detection while eliminating spurious duplicate alerts.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements feedback by continuously comparing new clusters with historical clusters and updating the cluster set accordingly. The comparison mechanism provides feedback on whether new abnormal terms represent new topics or variations of existing topics, enabling the system to learn from past patterns and reduce false alerts while maintaining accurate topic identification.

Inventive Principle:
Principle #23Feedback

2Object-generated harmful factors

If clustering algorithms are applied to group terms by topic, then some duplicate alerts are reduced, but multiple clusters still form for the same topic due to changing terminology over time

Engineering Contradiction:
Improveduplicate alertsVSAvoidcluster comparison mechanism
Core Design Contradiction:
Object-generated harmful factorsVSDevice complexity

Solution Approach 1:

The patent performs preliminary clustering within each time period to group abnormal terms into topic-related clusters before comparing clusters across time periods. This preliminary organization reduces the complexity of the subsequent comparison step by working with consolidated clusters rather than individual terms, while still achieving the goal of reducing duplicate alerts through temporal comparison.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple clusters are maintained for different time periods, then topic evolution is captured, but the system complexity increases

Engineering Contradiction:
Improvetopic evolution trackingVSAvoidcluster management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system discards clusters that are determined to represent the same topic as existing clusters, keeping only one representative cluster per topic. When terminology evolves and new clusters form, the system recovers by creating new clusters only when they represent genuinely new topics. This approach captures topic evolution while managing complexity by eliminating redundant clusters.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS11461406B2System and method for identifying newly trending topics in a data stream
Publication Date: 2022.10.04 VERINT SYST UK LTD
  • US11461406B2 patent drawing
  • US11461406B2 patent drawing
  • US11461406B2 patent drawing

AI summary

A system, computer implemented method, and computer storage medium encoded with a computer program, for identifying newly trending topics in a data stream. An example method includes: receiving text documents forming part of a data stream from one or more servers; identifying terms within the received text documents; deriving from the identified terms, a set of terms identified as abnormal by virtue of having a relatively high frequency of occurrence within the text documents received in a recent period compared with that expected from their historic occurrence; creating a first set of one or more clusters, each cluster including a group of terms from the set of terms identified as abnormal which through their degree of co-occurrence in the received text documents are considered to relate to the same topic; and comparing clusters of a further set with the clusters of the first set to determine whether a cluster of the further set pertains to the same topic.