Dataset Interrelation Discovery via Atomized Data Linking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage and computing technologies face challenges in managing and analyzing disparate datasets due to incompatibility among different platforms, systems, and formats, leading to isolated 'data silos' that hinder data interoperability and require significant manual effort for data scientists to assess and utilize datasets effectively.
Innovation Solution
A computerized system and interface that enables the discovery, formation, and analysis of collaborative datasets by linking atomized data points across disparate platforms, using programmatic interfaces and user interfaces to facilitate data interoperability, attribute derivation, and correlation, allowing for the creation of consolidated datasets with enhanced metadata and access control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional data storage and computing technologies are used to manage disparate datasets, then data can be stored in different platforms and formats, but data interoperability is hindered and data silos are created
Solution Approach 1:
The patent introduces a standardized data interface layer that acts as an intermediary between disparate data sources and analysis tools. This interface layer provides common protocols and formats that enable different platforms to communicate without requiring complex point-to-point integrations, thereby improving interoperability while managing system complexity.
Solution Approach 2:
The patent implements a universal data access framework that can work with multiple data formats and platforms through a single standardized interface. This multi-functional approach allows the same interface to handle various data types (structured, unstructured, semi-structured) across different platforms, enhancing adaptability without proportionally increasing complexity.
2Ease of operation
If manual assessment and analysis of datasets are performed by data scientists, then personalized queries and analyses can be conducted, but significant manual effort and time are required
Solution Approach 1:
The patent implements automated data profiling and quality assessment features that perform preliminary analysis of datasets before users engage with them. The system automatically generates metadata, identifies data quality issues, and suggests relevant analyses, thereby reducing the manual effort and time required for initial data assessment while maintaining ease of operation.
3Productivity
If contextual information is absent from ad hoc datasets, then data can be gathered quickly, but understanding and assessing dataset worthiness becomes complicated
Solution Approach 1:
The patent implements automated metadata generation and contextual information extraction that provides feedback about dataset quality, relevance, and usability. The system analyzes incoming data and automatically attaches contextual information, quality metrics, and relevance indicators, enabling users to quickly assess dataset value without sacrificing gathering speed.
4Adaptability or versatility
If different ad hoc approaches are used for gathering, forming, and analyzing datasets, then personalized data processing can be achieved, but inconsistency and incompatibility increase
Solution Approach 1:
The patent implements a hybrid approach where standardized protocols provide the foundation for consistent data processing, while allowing local customization and specialization at specific processing stages. This enables personalized data processing approaches to be applied locally without compromising overall system consistency and reliability.
Data Source
AI summary
Various techniques are disclosed for computerized tools to discover, form, and analyze dataset interrelations among a system of networked collaborative datasets including a repository configured to receive and store a dataset, and a dataset consolidation system configured to receive data to form a first input to initiate creation of a dataset based on a set of data, to activate a programmatic interface, to transform the set of data from a first format to an atomized format to form an atomized dataset, to monitor the creation of the dataset, to present data representing a status of a portion of the creation of the dataset, to calculate automatically dataset attributes of the linked dataset, to generate a plurality of sub-queries, and to retrieve data representing query results from the at least one of the different data repositories.


