Disambiguated Data Collections for Mining Interconnected Entities

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data mining systems are limited in analyzing datasets with ambiguities, as they struggle to disambiguate between different types of transactions or relationships, leading to unclear insights when analyzing interconnected data entities.

Innovation Solution

The method involves subdividing datasets into composite collections that are disambiguated for specific relationship scopes, allowing each collection to be analyzed separately by a data mining engine, thereby eliminating ambiguities and providing clear insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If data mining systems analyze individual data portions separately, then analysis simplicity is maintained, but the ability to analyze interconnected data entities together is lost

Engineering Contradiction:
Improveanalysis simplicityVSAvoidability to analyze interconnected data entities
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the dataset into multiple disambiguated collections, where each collection contains a subset of data entities and relationships that have been separated to eliminate ambiguities. This allows the data mining system to process each collection independently while still capturing interconnected relationships within each collection, thus maintaining analysis simplicity while improving adaptability to analyze interconnected entities.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If data mining systems attempt to analyze datasets with ambiguities, then comprehensive analysis coverage is achieved, but insight clarity deteriorates

Engineering Contradiction:
Improveanalysis coverageVSAvoidinsight clarity
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent extracts ambiguous relationships and data entities from the overall dataset and separates them into distinct disambiguated collections. By taking out the ambiguous elements and organizing them into separate collections with clear, unambiguous relationships, the system maintains comprehensive analysis coverage while ensuring that each collection provides clear, unambiguous insights suitable for data mining.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If data mining systems process the entire dataset as a single unit, then complete relationship analysis is possible, but processing efficiency decreases

Engineering Contradiction:
Improverelationship analysis completenessVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent divides the entire dataset into multiple smaller disambiguated collections that can be processed in parallel. Each collection maintains complete relationship analysis for its subset of entities, ensuring reliability, while the segmentation enables parallel processing and faster overall execution, thus improving productivity.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If data mining systems disambiguate datasets into multiple collections, then insight accuracy improves, but system complexity increases

Engineering Contradiction:
Improveinsight accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs disambiguation as a preliminary action before data mining processing. By pre-processing the dataset to create disambiguated collections with clear, unambiguous relationships, the system improves insight accuracy while reducing the complexity of the actual data mining process, as the collections are already prepared and organized for efficient processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10140344B2Extract metadata from datasets to mine data for insights
Publication Date: 2018.11.27 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10140344B2 patent drawing
  • US10140344B2 patent drawing
  • US10140344B2 patent drawing

AI summary

Analyzing data. A method includes obtaining a set of a plurality of data entities and relationships. The method further includes subdividing the set of a plurality of data entities and relationships into a plurality of composite collections of data entities and relationships. Each composite collection within the plurality of composite collections disambiguates the composite collection, within a relationship scope, from ambiguities in the set of a plurality of data entities and relationships. The method further includes providing one or more of the plurality of composite collections of data entities and relationships to a data mining engine. Each composite collection of data entities and relationships is provided as a separate unit to the data mining engine.