Atomized Dataset Loading for Cross-Platform Data Store Queries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage and access technologies are inadequate for managing and interoperating large, disparate datasets due to format incompatibilities and resource limitations, leading to data silos that hinder collaboration and knowledge sharing among researchers and organizations.
Innovation Solution
A collaborative dataset consolidation system that converts datasets into atomized form, enabling interoperability and secure access across platforms, using a dataset ingestion controller to link and consolidate datasets, and a collaboration manager to facilitate data sharing and correlation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If datasets are stored in conventional data stores with different formats and platforms, then data storage capacity is increased, but data interoperability and accessibility are reduced
Solution Approach 1:
The patent segments datasets into atomic data points that can be independently stored, managed, and accessed. Each atomized data point represents a discrete unit of information that can be combined with other atomized points from different sources, enabling interoperability while maintaining storage capacity for large volumes of diverse data.
Solution Approach 2:
The patent introduces an intermediary layer (the atomization framework and collaboration manager) that translates between different data formats and platforms. This intermediary enables datasets from diverse sources to be converted into a common atomized representation, facilitating interoperability without losing the ability to store large quantities of varied data.
2Ease of operation
If datasets are consolidated into a unified system, then data accessibility and collaboration are improved, but system complexity and resource requirements increase
Solution Approach 1:
By segmenting the consolidation process into atomic data points, the system avoids the complexity of managing entire heterogeneous datasets as unified entities. Each atomized point can be independently processed, stored, and accessed, simplifying the overall system architecture while improving accessibility.
Solution Approach 2:
The patent changes the fundamental parameter of data representation from format-specific structures to a universal atomized format. This parameter change enables simplified access and collaboration mechanisms while the system manages the complexity of handling diverse source data through standardized transformation processes.
3Speed
If atomized data points are stored in memory, then data access speed is increased, but memory resource consumption increases
Solution Approach 1:
The patent segments data into atomic points that can be selectively loaded into memory based on query requirements. Instead of loading entire datasets, only relevant atomized points are retrieved and stored in memory, enabling fast access while minimizing memory consumption through on-demand loading of specific data elements.
Data Source
AI summary
Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a system may include an atomized workflow loader configured to receive an atomized dataset to load into a data store, and to determine resource requirements data to describe at least one resource requirement. The atomized workflow loader may be further configured to select a data store type based on a resource requirement, and perform a load operation of the atomized dataset as a function of the data store type.


