Hierarchical Clustering for Enterprise Data Records

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managing large volumes of unstructured text data in information processing systems is tedious and time-consuming due to the need for manual screening and customization of rules to define relationships between data records in graph networks.

Innovation Solution

The implementation of hierarchical clustering techniques, which involve generating similarity matrices, applying thresholding filters to create adjacency matrices, and constructing graph networks to identify clusters in data records, enabling automated remedial actions within enterprise systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual screening and customization of rules are used to define relationships between data records, then relationships can be accurately defined, but the process becomes unduly tedious and time-consuming

Engineering Contradiction:
Improveaccuracy of relationship definitionVSAvoidtime-consuming processing
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical screening and rule customization with automated computational methods. Specifically, it uses graph neural networks to automatically learn relationships between data records from unstructured text data, eliminating the need for manual rule definition while maintaining relationship accuracy. The system automatically constructs graph networks where nodes represent data records and edges represent learned relationships, achieving both precision and efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the graph neural network to autonomously learn and define relationships between data records without human intervention. The network automatically processes unstructured text, identifies patterns, and constructs the graph structure itself, making the system self-sufficient in defining relationships that would otherwise require manual curation.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If manual rules are customized to process unstructured text data, then specific themes can be identified, but the process becomes tedious and requires large volumes of manual effort

Engineering Contradiction:
Improveability to identify themesVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent replaces manual theme identification rules with a graph neural network that automatically learns thematic patterns from unstructured text data. The network processes text records, extracts relevant features, and identifies themes through learned representations in the graph structure, eliminating tedious manual rule customization while maintaining thematic identification capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the approach from fixed manual rules to dynamic learned parameters. The graph neural network learns optimal parameters for theme identification directly from the data, allowing the system to adapt to different themes and domains without requiring manual rule reconfiguration. This parameter learning approach maintains versatility while dramatically improving processing efficiency.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If graph networks are constructed with explicit relationships defined manually, then contextual information can be provided, but the construction process becomes time-consuming

Engineering Contradiction:
Improvecompositional or contextual informationVSAvoidgraph construction time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent replaces manual graph construction with automated graph neural network processing. The network automatically constructs graph networks from unstructured text data, where nodes represent data records and edges represent learned contextual relationships. This automated construction preserves compositional and contextual information while reducing graph construction time from manual processes to computational operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary action by pre-processing unstructured text data into structured graph representations that capture contextual information. The graph neural network learns and stores relationships in the graph structure during the construction phase, enabling efficient subsequent queries and analysis without requiring manual relationship definition each time the graph is needed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11599568B2Monitoring an enterprise system utilizing hierarchical clustering of strings in data records
Publication Date: 2023.03.07 EMC IP HLDG CO LLC
  • US11599568B2 patent drawing
  • US11599568B2 patent drawing
  • US11599568B2 patent drawing

AI summary

An apparatus includes a processing device configured to obtain data records associated with an enterprise system comprising strings associated with an attribute. The processing device is also configured to generate a similarity matrix with entries comprising values characterizing similarity between respective pairs of the strings. The processing device is further configured to apply a thresholding filter to values in the entries of the similarity matrix to create an adjacency matrix, and to construct a graph network of the data records based at least in part on the adjacency matrix, wherein the graph network comprises edges connecting pairs of the data records. The processing device is further configured to perform a clustering operation on the graph network to identify clusters of the data records for the attribute, and to initiate remedial action in the enterprise system responsive to identifying a given cluster comprising a given subset of the data records.