Collaborative Dataset Consolidation With Atomized Data Linking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage and access technologies are inadequate for managing and interoperating large, disparate datasets due to format incompatibilities and resource limitations, leading to data silos that hinder collaboration and knowledge sharing among researchers and organizations.
Innovation Solution
A collaborative dataset consolidation system that converts datasets into atomized form, allowing linking and interoperability across different formats, with granular security and authorization controls, and facilitates correlations and collaborations among users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If datasets are stored in conventional data stores as data silos to preserve format compatibility and security, then data security and format integrity are maintained, but data interoperability and collaboration efficiency deteriorate
Solution Approach 1:
The patent introduces an intermediary layer (atomization service, graph database) that mediates between disparate data formats and storage locations. This intermediary converts diverse datasets into a common atomized representation, enabling interoperability without requiring changes to source systems or compromising security boundaries.
Solution Approach 2:
The patent segments datasets into atomic units (atomized datasets) that can be independently stored, managed, and linked. This segmentation allows different datasets to maintain their original security and format characteristics while being connectable through standardized links in the graph database.
2Productivity
If datasets are consolidated into a unified system to improve collaboration and access efficiency, then data accessibility and collaboration improve, but system complexity and resource requirements increase
Solution Approach 1:
The patent adds a new dimensional layer (graph database layer) above existing data storage systems. This additional dimension enables complex relationships and collaborations to be represented and queried without modifying the underlying storage systems, thereby improving access efficiency while containing complexity in a separate layer.
Solution Approach 2:
The patent creates a universal atomized data format and linkage mechanism that can represent and connect any type of dataset from any source. This universal approach consolidates multiple data access pathways into a single system that handles diverse data types, reducing overall system complexity despite the variety of sources.
3Adaptability or versatility
If datasets are atomized and linked across multiple repositories to enhance interoperability, then data collaboration and discovery improve, but data security and access control become more challenging
Solution Approach 1:
The patent segments access control into two distinct layers: repository-level security (managing physical data storage and access) and link-level security (managing relationships between atomized datasets). This segmentation allows collaboration to proceed through public links while sensitive operations require authentication, reducing overall access control complexity.
Solution Approach 2:
The patent introduces an intermediary authentication and authorization layer that mediates access requests between users and atomized datasets. This intermediary handles security credentials and permissions centrally, enabling fine-grained access control across distributed repositories without requiring complex security implementations at each storage location.
4Quantity of substance
If conventional data storage technologies are used to manage vast amounts of diverse data, then data storage capacity is sufficient, but data analysis and interoperability capabilities are inadequate
Solution Approach 1:
The patent changes the fundamental parameter of data representation from conventional structured formats (tables, documents) to atomized graph-based representations. This parameter change enables the same storage capacity to support sophisticated analysis and interoperability operations that are not possible with traditional data formats.
Solution Approach 2:
The patent introduces a graph database intermediary that sits between conventional storage systems and analysis operations. This intermediary translates various data formats into a unified graph structure, enabling powerful analysis and interoperability capabilities while preserving the underlying storage infrastructure.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby a collaborative data layer and associated logic facilitate, for example, efficient access to, and implementation of, collaborative datasets. In some examples, a method may include receiving data representing a query into a collaborative dataset consolidation system, identifying datasets relevant to the query, generating one or more queries to access disparate data repositories, and retrieving data representing query results. In some cases, one or more queries are applied (e.g., as a federated query) to atomized datasets stored in one or more atomized data stores, at least two of which may be different.