Extensible Data Enclave for Variable Schema Ingestion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data warehousing systems struggle with variable data structures and mutating data attributes, leading to non-functional data lakes and broken data analysis due to the fast pace of software development and changing data formats.
Innovation Solution
A platform that forms an extensible data warehouse with a data ingestor application to manage and process data with flexible scalability, automatically detect data formats, and provide rigorous data security controls, enabling continuous data ingestion and transformation with minimal human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fixed data structures are used for data storage and access, then data access control and security are improved, but the system cannot accommodate variable data structures and mutating data attributes from diverse data sources
Solution Approach 1:
The patent introduces a data ingestor application as an intermediary component between diverse data sources and the data lake. This mediator automatically detects data formats, validates data structures, and transforms variable data into a standardized internal representation, thereby maintaining both security through controlled access and adaptability through automatic format recognition
Solution Approach 2:
The system dynamically changes data representation parameters by detecting and adapting to different data formats (JSON, CSV, XML, etc.) and structures. The data ingestor modifies data parameters on-the-fly to match expected schemas, enabling the system to handle mutating data attributes while maintaining consistent internal storage structures for security
2Quantity of substance
If data is continuously ingested from multiple sources with varying structures, then data volume and analytics capability are improved, but data analysis and big data pipeline tools become nonfunctional due to structure changes
Solution Approach 1:
The data ingestor performs preliminary actions by detecting data formats and validating structures before data enters the lake. It proactively transforms and normalizes data in advance, preventing structure changes from propagating through the pipeline and causing tool failures
Solution Approach 2:
The system implements feedback mechanisms where the data ingestor continuously monitors data structures, compares them against expected schemas, and adjusts processing behavior accordingly. This feedback loop ensures data analysis tools receive consistent, validated data regardless of source variations
3Adaptability or versatility
If manual processing of variable data structures is performed, then data can be accommodated, but human intervention is required for each data structure change
Solution Approach 1:
The data ingestor application provides self-service by automatically detecting data formats, inferring schemas, and transforming variable data structures without human intervention. The system serves itself by maintaining metadata catalogs that enable autonomous adaptation to new data formats
Solution Approach 2:
The system creates copies of data in standardized internal representations rather than storing raw variable-format data. This copying approach allows the system to preserve the original variable structures for reference while working with normalized versions, enabling automated processing without manual intervention
Data Source
AI summary
Systems, methods, and non-transitory computer-readable media for forming an extensible data warehouse. A data ingestor application receiving raw data having a first structure. Forming a data lake in the third memory using the raw data. Continuously receive additional raw data having a plurality of structures. The plurality of structures including the first structure and one or more different structures. The additional raw data supplementing the raw data. Determining each structure of the plurality of structures. Generating a dataset based on the additional raw data and the plurality of structures that are determined. Extracting metadata associated with the additional raw data from the dataset. Creating a catalog of the dataset based on the metadata that is extracted. Modifying the data lake in the third memory to include the additional raw data based on the dataset and the catalog.


