Co-clustering Strings in Data Records for Graph Network Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing large volumes of unstructured text data in information processing systems is tedious and time-consuming due to the need for manual screening and customization of rules to define relationships between data records in graph networks.
Innovation Solution
The technique involves obtaining sets of data records with associated strings, generating a similarity matrix to characterize similarity between strings, constructing a graph network based on this matrix, performing co-clustering operations to identify clusters, and initiating remedial actions in response to identified clusters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual screening and customization of rules are used to define relationships between data records, then the relationships can be accurately defined, but the process becomes tedious and time-consuming
Solution Approach 1:
The patent replaces manual mechanical screening and rule customization with automated computational methods. Specifically, it uses graph network construction where data records become nodes and relationships are automatically determined through algorithmic processing of unstructured text data, eliminating the need for manual rule definition while maintaining relationship accuracy
Solution Approach 2:
The system enables self-service by allowing the graph network construction process to automatically identify and define relationships between data records without human intervention. The automated clustering algorithms and similarity computations perform the relationship definition task independently, freeing operators from tedious manual work
2Measurement precision
If manual processing methods are used for unstructured text data, then detailed analysis can be performed, but the productivity decreases for large volumes of data
Solution Approach 1:
The patent substitutes manual text analysis with automated computational approaches including graph network construction and clustering algorithms. These systems process unstructured text data at machine speed while maintaining analytical depth, enabling detailed relationship identification across large datasets that would be impossible to analyze manually
Solution Approach 2:
The patent segments the complex task of unstructured text analysis into distinct computational components: graph network construction, similarity matrix generation, and clustering operations. This segmentation allows each component to be optimized independently and processed in parallel, dramatically increasing productivity while preserving analytical detail
3Measurement precision
If custom rules are created for each data theme, then specific themes can be accurately identified, but the device complexity increases due to maintenance of large rule sets
Solution Approach 1:
The patent replaces the complex mechanical system of custom rule creation and maintenance with automated graph-based relationship identification. The system dynamically determines relationships between data records through computational analysis of unstructured text, eliminating the need for pre-defined rules for each theme while maintaining accurate theme identification
Solution Approach 2:
The patent implements a universal graph network construction approach that can identify multiple different themes and relationships simultaneously without requiring separate custom rules for each. The same automated system handles diverse data types and relationships, reducing complexity while maintaining thematic accuracy
Data Source
AI summary
An apparatus includes a processing device configured to obtain first and second sets of data records, each data record comprising a string associated with an attribute. The processing device is also configured to generate a similarity matrix, wherein entries of the similarity matrix comprise values characterizing similarity between respective pairs of the strings comprising a first string from a data record in the first set and a second string from a data record in the second set. The processing device is further configured to construct a graph network based on the similarity matrix comprising edges connecting pairs of the data records based on values of entries in the similarity matrix, perform a clustering operation on the graph network to identify clusters, and to initiate remedial action responsive to identifying a given cluster comprising at least one data record from each of the first and second sets of data records.


