Disambiguated Data Collections for Mining Interconnected Entities
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data mining systems are limited in analyzing datasets with ambiguities, as they struggle to disambiguate between different types of transactions or relationships, leading to unclear insights when analyzing interconnected data entities.
Innovation Solution
The method involves subdividing datasets into composite collections that are disambiguated for specific relationship scopes, allowing each collection to be analyzed separately by a data mining engine, thereby eliminating ambiguities and providing clear insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If data mining systems analyze individual data portions separately, then analysis simplicity is maintained, but the ability to analyze interconnected data entities together is lost
Solution Approach 1:
The patent segments the dataset into multiple disambiguated collections, where each collection contains a subset of data entities and relationships that have been separated to eliminate ambiguities. This allows the data mining system to process each collection independently while still capturing interconnected relationships within each collection, thus maintaining analysis simplicity while improving adaptability to analyze interconnected entities.
2Quantity of substance
If data mining systems attempt to analyze datasets with ambiguities, then comprehensive analysis coverage is achieved, but insight clarity deteriorates
Solution Approach 1:
The patent extracts ambiguous relationships and data entities from the overall dataset and separates them into distinct disambiguated collections. By taking out the ambiguous elements and organizing them into separate collections with clear, unambiguous relationships, the system maintains comprehensive analysis coverage while ensuring that each collection provides clear, unambiguous insights suitable for data mining.
3Reliability
If data mining systems process the entire dataset as a single unit, then complete relationship analysis is possible, but processing efficiency decreases
Solution Approach 1:
The patent divides the entire dataset into multiple smaller disambiguated collections that can be processed in parallel. Each collection maintains complete relationship analysis for its subset of entities, ensuring reliability, while the segmentation enables parallel processing and faster overall execution, thus improving productivity.
4Measurement precision
If data mining systems disambiguate datasets into multiple collections, then insight accuracy improves, but system complexity increases
Solution Approach 1:
The patent performs disambiguation as a preliminary action before data mining processing. By pre-processing the dataset to create disambiguated collections with clear, unambiguous relationships, the system improves insight accuracy while reducing the complexity of the actual data mining process, as the collections are already prepared and organized for efficient processing.
Data Source
AI summary
Analyzing data. A method includes obtaining a set of a plurality of data entities and relationships. The method further includes subdividing the set of a plurality of data entities and relationships into a plurality of composite collections of data entities and relationships. Each composite collection within the plurality of composite collections disambiguates the composite collection, within a relationship scope, from ambiguities in the set of a plurality of data entities and relationships. The method further includes providing one or more of the plurality of composite collections of data entities and relationships to a data mining engine. Each composite collection of data entities and relationships is provided as a separate unit to the data mining engine.


