Dataset Ingestion Controller for Interoperable Data Consolidation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage and computing technologies are inadequate for managing and accessing vast, disparate datasets due to incompatibility among different platforms and formats, leading to isolated 'data silos' that hinder data interoperability and limit access to valuable information, especially for resource-constrained organizations and researchers.
Innovation Solution
A collaborative dataset consolidation system that converts datasets into a unified, atomized format, enabling interoperability across disparate platforms and formats, and facilitates secure, authorized access and sharing through a dataset ingestion controller, query engine, and collaboration manager.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If datasets are stored in conventional data stores as separate entities, then data storage capacity is maintained, but data interoperability and accessibility are hindered due to format incompatibility and platform differences
Solution Approach 1:
The patent segments datasets into standardized, interoperable formats that can be independently accessed and combined. By breaking down complex data structures into manageable, standardized components, the system enables different datasets to work together across platforms without requiring complete system redesign.
Solution Approach 2:
The patent implements a universal data access interface that can handle multiple data formats and platforms through a single standardized protocol. This multi-functional approach allows the same system to interact with diverse data sources (scientific, corporate, academic) without requiring format-specific handling, thereby improving interoperability while managing complexity.
2Ease of operation
If data is consolidated into a unified system, then data accessibility and collaboration are improved, but security risks and authorization management complexity increase
Solution Approach 1:
The patent introduces an intermediary layer (data access interface and collaboration manager) that sits between users and the consolidated datasets. This intermediary handles authorization, authentication, and access control policies, thereby maintaining security while enabling efficient data access. The intermediary translates user requests into authorized operations without exposing the underlying data structure.
Solution Approach 2:
The system implements self-service authorization mechanisms where data owners can define access policies and users can self-register and request access according to predefined rules. This automated authorization management reduces the burden on central administrators while maintaining security controls, enabling efficient access to consolidated datasets without compromising reliability.
3Quantity of substance
If resource-constrained organizations access centralized data repositories, then data availability improves, but network dependency and access latency increase
Solution Approach 1:
The patent implements preliminary data processing and caching mechanisms where frequently accessed data is pre-loaded into local or edge caches. This preliminary action reduces the need for repeated network transactions, thereby maintaining high data availability while minimizing access latency for resource-constrained organizations.
Solution Approach 2:
The patent introduces a multi-dimensional data access architecture that combines centralized repository access with distributed caching layers. By adding this temporal and spatial dimension to data access, the system maintains data availability from the central repository while providing low-latency access through local caches, effectively resolving the network dependency issue.
Data Source
AI summary
Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving a dataset having a data format into a dataset ingestion controller configured to form a collaborative dataset, interpreting data of the dataset against data classifications at an inference engine to derive at least an inferred attribute, associating the data with annotative data identifying the inferred attribute, and converting the dataset at a format converter to form an atomized dataset.


