Dataset Similarity Determination for Tabular and Graph Data Interoperability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional technologies struggle to efficiently match and join datasets of different formats and structures, such as tabular and graph-based data, due to limitations in data interoperability and the complexity of large-scale data arrangements.
Innovation Solution
A collaborative dataset consolidation system that determines degrees of similarity between subsets of tabular data and graph-based data arrangements using similarity matrices and compressed data representations, facilitating the identification of relevant datasets for joining and enhancing data interoperability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional data storage and computing technologies are used to store and access large-scale datasets, then data storage capacity is sufficient, but data accessing and analysis efficiency deteriorates due to the complexity and size of datasets
Solution Approach 1:
The patent segments large-scale datasets into smaller manageable units organized in graph-based data arrangements. By dividing the data into interconnected nodes and edges representing entities and relationships, the system enables efficient traversal and analysis of specific data portions without processing the entire dataset, thus maintaining storage capacity while improving access efficiency.
Solution Approach 2:
The patent transitions from traditional tabular data arrangements to graph-based data arrangements, adding a dimensional shift in data organization. This graph structure introduces new dimensions for data access through nodes, edges, and paths, enabling multi-dimensional queries and analyses that improve efficiency compared to conventional row-column based approaches.
2Adaptability or versatility
If different data formats and structures (tabular and graph-based) are used to represent datasets, then data representation flexibility is improved, but data interoperability and matching complexity increases
Solution Approach 1:
The patent implements a universal graph-based data arrangement framework that can represent both traditional tabular data and complex relational data in a unified structure. This multi-functional approach allows the same graph infrastructure to handle diverse data types and formats, reducing the complexity of data interoperability by providing a common language for data representation across different formats.
Solution Approach 2:
The patent introduces similarity matrices as intermediary structures that facilitate matching between tabular data subsets and graph-based data subsets. These matrices act as mediators that translate and compare different data representations, enabling interoperability between tabular and graph formats without requiring complex direct mapping logic, thus reducing matching complexity while maintaining flexibility.
3Reliability
If manual intervention is used to identify and join relevant datasets, then joining accuracy is maintained, but processing time and operational effort increase
Solution Approach 1:
The patent implements automated similarity determination mechanisms that enable datasets to self-identify relevant matches through computational comparison of their structural and semantic characteristics. The system automatically calculates similarity metrics between data subsets, ranks potential matches, and suggests join operations without requiring manual intervention, thus maintaining high joining accuracy through algorithmic precision while dramatically reducing processing time and operational effort.
Solution Approach 2:
The patent incorporates feedback loops where the results of similarity comparisons and dataset joins are continuously evaluated and used to refine future matching operations. The system learns from previous joining operations, adjusting similarity thresholds and weighting schemes to improve accuracy over time, while the automated feedback mechanism eliminates the need for manual review, reducing operational effort while maintaining or improving joining accuracy.
Data Source
AI summary
Various techniques are described, including evaluating ingested data including a dataset to identify one or more links to other datasets stored in a graph, using a similarity determination algorithm to identify a degree of similarity between datasets to determine joinability of ingested datasets with graph-stored datasets, determining a ratio to determine whether to perform an overlap or coverage function, associating a subset of similarity matrices with a subset of graph data joined to the ingested dataset, and forming links in a column of data between the dataset and the another dataset of the ingested data based on the degree of similarity.


