Collaborative Dataset Consolidation via Inference Engine Anomaly Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data management techniques face challenges in efficiently consolidating and interoperating disparate datasets due to incompatibilities in data formats, storage systems, and manual intervention requirements, leading to inefficiencies in data analysis and sharing.
Innovation Solution
A collaborative dataset consolidation system that converts datasets into a unified format, using a dataset ingestion controller with an inference engine and layer data generator to create layered data files, enabling data interoperability and automatic anomaly detection and remediation, facilitating the formation of interrelated datasets across different platforms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional data management techniques are used to consolidate disparate datasets, then data from different sources can be stored, but data interoperability and compatibility between different formats and systems deteriorate
Solution Approach 1:
The patent introduces a standardized data format and conversion layer as an intermediary between disparate data sources and the consolidation system. This intermediary enables different data formats (CSV, TSV, HTML, JSON, XML) to be transformed into a unified structure, allowing data from diverse sources to be consolidated while maintaining interoperability through standardization protocols.
2Reliability
If manual intervention is used to standardize data arrangements and group data, then data quality and consistency improve, but operational efficiency and productivity deteriorate
Solution Approach 1:
The system implements automated data standardization through inference engines that automatically detect data types, group data by attributes, and standardize arrangements without manual intervention. The system self-services by autonomously performing tasks that traditionally required data practitioners, such as deciding how to group data and standardizing formats, thereby maintaining consistency while dramatically improving productivity.
Solution Approach 2:
The patent replaces manual mechanical processes of data standardization with automated computational systems. Inference engines and algorithms substitute for human data practitioners, automatically analyzing data characteristics, determining appropriate groupings, and applying standardization rules, thus eliminating the friction of manual intervention while preserving data quality.
3Quantity of substance
If datasets are stored in conventional data silos with different computing platforms and database technologies, then data storage capacity increases, but data access and analysis interoperability deteriorate
Solution Approach 1:
The patent creates a universal data access interface that can interact with multiple different computing platforms and database technologies through a single standardized protocol. The system performs multiple functions by supporting various data formats and storage systems while presenting a unified access method, enabling seamless data retrieval and analysis across diverse platforms without requiring separate access mechanisms for each system.
4Productivity
If free-form data formats are imported without manual intervention, then data import speed increases, but data structure consistency and reliability deteriorate
Solution Approach 1:
The system performs preliminary automated standardization actions during the data import process itself. Rather than importing raw free-form data and requiring subsequent manual correction, the inference engine proactively analyzes incoming data, determines appropriate structures, and applies standardization transformations in advance, ensuring consistency is established before data is fully integrated into the system.
Data Source
AI summary
Various embodiments relate generally to data science and data analysis, computer software and systems, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby logic is configured to remediate anomalies in a data set originating in a first format prior to enrichment and conversion into a second format that facilitates forming collaborative dataset and, for example, interrelations among a system of networked collaborative datasets, whereby, at least in some implementations, data interrelations between different formats may be disposed in one or more data layers (e.g., layered data files and/or data arrangements). In some examples, a method may converting a dataset from a data format at a format converter to form an atomized dataset in a graph data arrangement, the atomized dataset being a collaborative dataset including atomized descriptor data and atomized source data.


