Collaborative Dataset Consolidation via Layered Data Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data management techniques are inadequate for efficiently consolidating and analyzing datasets across disparate platforms, leading to data silos that hinder interoperability, require manual intervention for standardization, and are inefficient in detecting anomalies, thereby reducing the reliability and usability of datasets.
Innovation Solution
A collaborative dataset consolidation system that converts datasets into a unified format, using a dataset ingestion controller with an inference engine and layer data generator to form interrelated data layers, enabling interoperability and anomaly detection across different data formats and platforms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional data management techniques are used to consolidate datasets across disparate platforms, then data interoperability is improved, but manual intervention is still required for standardization and anomaly detection
Solution Approach 1:
The system enables datasets to self-describe their schemas, formats, and relationships through automated metadata generation and inference engines that automatically detect anomalies and standardize data without human intervention, allowing the data infrastructure to service itself
Solution Approach 2:
Manual mechanical processes of data standardization and anomaly detection are replaced with automated computational systems including inference engines, schema validators, and machine learning models that perform these functions algorithmically
2Quantity of substance
If conventional data storage technologies are used to store increasing amounts of data, then data capacity is improved, but data accessibility and analysis efficiency deteriorate due to data silos
Solution Approach 1:
The system creates a universal data access layer that can read and interpret multiple data formats and schemas simultaneously, allowing a single infrastructure to serve multiple data sources and analysis purposes without creating silos
Solution Approach 2:
An intermediary data access layer and metadata system is introduced between raw data storage and analysis tools, enabling efficient querying and analysis across distributed datasets without requiring direct access to each individual data source
3Reliability
If manual data standardization is performed to ensure data quality, then data reliability is improved, but time consumption and operational complexity increase
Solution Approach 1:
Data standardization rules, schemas, and validation criteria are pre-configured and automatically applied to incoming datasets, performing standardization actions in advance before analysis occurs, eliminating the need for time-consuming manual standardization processes
4Difficulty of detecting and measuring
If conventional anomaly detection methods are used to identify data issues, then detection capability is improved, but false positives and detection accuracy remain problematic
Solution Approach 1:
The system implements feedback loops where detection results are continuously refined based on validation outcomes and user corrections, with the inference engine learning from detected patterns to improve future anomaly detection accuracy and reduce false positives
Data Source
AI summary
Various embodiments relate generally to data science and data analysis, computer software and systems, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby logic is configured to remediate anomalies in a data set originating in a first format prior to enrichment and conversion into a second format that facilitates forming collaborative dataset and, for example, interrelations among a system of networked collaborative datasets, whereby, at least in some implementations, data interrelations between different formats may be disposed in one or more data layers (e.g., layered data files and/or data arrangements). In some examples, a method may include analyzing data to detect a non-compliant data attribute, detecting a condition based on the non-compliant data attribute, invoking an action to modify a subset of data, and generating a graph data arrangement linkable to other graph data arrangements to form a collaborative dataset.


