Metadata-Driven Data Lake Ingestion Engine
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management systems face delays and inconsistencies in ingesting data into Enterprise Data Lakes due to manual processes and the need for IT intervention, lacking self-serve capabilities for data providers and consumers, and inadequate security control over datasets.
Innovation Solution
A system with an ingestion engine and metadata model that automatically determines metadata type, generates processing guidelines, transforms data, and applies security policies, enabling self-serve data ingestion, standardization, and security management within a data lake.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual ETL processes and SDLC are used for data ingestion, then data can be ingested into the data lake, but lengthy delays occur between data request and availability
Solution Approach 1:
The system enables self-serve data ingestion by allowing data providers to upload datasets with embedded metadata that automatically drives the ingestion process, eliminating the need for manual ETL development and SDLC approval cycles. The metadata model with executable instructions allows the system to autonomously process data without human intervention.
Solution Approach 2:
Data providers perform preliminary actions by embedding complete metadata instructions with their datasets before upload. This metadata includes all necessary transformation rules, target schema definitions, and processing guidelines that enable the system to execute the entire ingestion workflow immediately upon upload, bypassing traditional development and approval timelines.
2Manufacturing precision
If manual ETL development is required, then data can be standardized, but inconsistencies arise due to different developers' approaches
Solution Approach 1:
The system changes the fundamental parameter of data standardization from developer-dependent manual processes to metadata-driven automated processes. By encoding transformation rules, target schemas, and processing parameters directly in the metadata model, the system ensures consistent application of standardization rules regardless of which data provider or consumer initiates the process.
Solution Approach 2:
The metadata model acts as an intermediary between data providers and the ingestion system. Instead of developers manually creating ETL code, the metadata serves as a structured intermediary that translates data provider requirements into standardized ingestion instructions, eliminating variability introduced by different developers' approaches.
3Reliability
If IT intervention is required for data ingestion, then data security can be controlled, but self-serve capabilities are lost
Solution Approach 1:
The system implements self-serve data ingestion with automated security control by embedding security policies and access control instructions directly in the metadata model. Data providers can specify security requirements, access permissions, and compliance rules in their metadata, which the system automatically enforces during ingestion and storage, eliminating the need for manual IT security reviews while maintaining strong security controls.
Data Source
AI summary
The present invention relates, in an embodiment, to a system for automatically ingesting data into a data lake. In an embodiment of the present invention, the system comprises computer readable memory having recorded thereon instructions for execution by a processor having an ingestion engine and a metadata model. In an embodiment of the present invention, the instructions are configured to determine, via the metadata model, a type of metadata the ingested data contains; to generate guidelines for processing and transforming the ingested data based on the determined metadata; to apply the guidelines at the ingestion engine for how the ingestion engine processes and transforms the ingested data based on the determined metadata; and to store the transformed ingested data to a storage repository.


