Hierarchical Clustering for Enterprise Data Records
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing large volumes of unstructured text data in information processing systems is tedious and time-consuming due to the need for manual screening and customization of rules to define relationships between data records in graph networks.
Innovation Solution
The implementation of hierarchical clustering techniques, which involve generating similarity matrices, applying thresholding filters to create adjacency matrices, and constructing graph networks to identify clusters in data records, enabling automated remedial actions within enterprise systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual screening and customization of rules are used to define relationships between data records, then relationships can be accurately defined, but the process becomes unduly tedious and time-consuming
Solution Approach 1:
The patent replaces manual mechanical screening and rule customization with automated computational methods. Specifically, it uses graph neural networks to automatically learn relationships between data records from unstructured text data, eliminating the need for manual rule definition while maintaining relationship accuracy. The system automatically constructs graph networks where nodes represent data records and edges represent learned relationships, achieving both precision and efficiency.
Solution Approach 2:
The system enables self-service by allowing the graph neural network to autonomously learn and define relationships between data records without human intervention. The network automatically processes unstructured text, identifies patterns, and constructs the graph structure itself, making the system self-sufficient in defining relationships that would otherwise require manual curation.
2Adaptability or versatility
If manual rules are customized to process unstructured text data, then specific themes can be identified, but the process becomes tedious and requires large volumes of manual effort
Solution Approach 1:
The patent replaces manual theme identification rules with a graph neural network that automatically learns thematic patterns from unstructured text data. The network processes text records, extracts relevant features, and identifies themes through learned representations in the graph structure, eliminating tedious manual rule customization while maintaining thematic identification capability.
Solution Approach 2:
The system changes the approach from fixed manual rules to dynamic learned parameters. The graph neural network learns optimal parameters for theme identification directly from the data, allowing the system to adapt to different themes and domains without requiring manual rule reconfiguration. This parameter learning approach maintains versatility while dramatically improving processing efficiency.
3Loss of information
If graph networks are constructed with explicit relationships defined manually, then contextual information can be provided, but the construction process becomes time-consuming
Solution Approach 1:
The patent replaces manual graph construction with automated graph neural network processing. The network automatically constructs graph networks from unstructured text data, where nodes represent data records and edges represent learned contextual relationships. This automated construction preserves compositional and contextual information while reducing graph construction time from manual processes to computational operations.
Solution Approach 2:
The system performs preliminary action by pre-processing unstructured text data into structured graph representations that capture contextual information. The graph neural network learns and stores relationships in the graph structure during the construction phase, enabling efficient subsequent queries and analysis without requiring manual relationship definition each time the graph is needed.
Data Source
AI summary
An apparatus includes a processing device configured to obtain data records associated with an enterprise system comprising strings associated with an attribute. The processing device is also configured to generate a similarity matrix with entries comprising values characterizing similarity between respective pairs of the strings. The processing device is further configured to apply a thresholding filter to values in the entries of the similarity matrix to create an adjacency matrix, and to construct a graph network of the data records based at least in part on the adjacency matrix, wherein the graph network comprises edges connecting pairs of the data records. The processing device is further configured to perform a clustering operation on the graph network to identify clusters of the data records for the attribute, and to initiate remedial action in the enterprise system responsive to identifying a given cluster comprising a given subset of the data records.


