Dynamic Data-Ingestion Pipeline for Hadoop
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Handling high volumes of heterogeneous data from various sources poses challenges in computer systems, including increased complexity, high costs, and degraded performance due to the need for significant memory and offline processing.
Innovation Solution
A dynamic data-ingestion pipeline is generated using a modular arrangement of operators, allowing real-time processing without storing data in memory, and adaptable to various data sources, including Hadoop file systems, to ingest and transform data efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If significant memory is used to store large amounts of data from different data sources, then data storage capacity is improved, but system cost and complexity increase
Solution Approach 1:
The patent extracts only the necessary data fields from incoming data streams rather than storing entire records. The schema-based extraction mechanism identifies and pulls only relevant fields, significantly reducing memory requirements while maintaining data utility for processing
Solution Approach 2:
The data processing system is segmented into multiple independent components: data ingestion services, schema validation modules, extraction operators, and output writers. Each component handles specific tasks independently, reducing overall system complexity while maintaining high data storage capacity
2Loss of energy
If offline or non-real-time processing is used, then system cost is reduced, but performance and user experience degrade
Solution Approach 1:
The patent implements dynamic data extraction where the system adapts its processing behavior based on incoming data characteristics and schema definitions. The extraction operators dynamically adjust which fields to pull based on schema validation results, enabling real-time processing without static rigid configurations
Solution Approach 2:
A schema validation intermediary layer is introduced between data ingestion and extraction. This intermediary validates incoming data against defined schemas and guides the extraction process, enabling real-time processing decisions without requiring complete offline analysis
3Adaptability or versatility
If separate data processing is implemented for each data source, then data source compatibility is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal data extraction framework that handles multiple data sources through a common interface. The schema-based extraction operators work uniformly across different data sources (Hadoop, relational databases, NoSQL), providing multi-functionality without requiring separate processing logic for each source
Solution Approach 2:
The system uses schema copying and template matching to handle different data sources. Instead of creating custom processing logic for each source, the system copies and adapts schema templates to match various data source structures, enabling compatibility through pattern replication rather than custom implementation
4Speed
If high velocity data integration is implemented, then data processing speed is improved, but system cost increases
Solution Approach 1:
The patent applies partial action by extracting only the necessary subset of fields from incoming data rather than processing and storing all data. This selective extraction achieves high velocity integration for the critical data fields without the cost of processing entire data records
Data Source
AI summary
In order to ingest data from an arbitrary source in a set of sources, a computer system accesses predefined configuration instructions. Then, the computer system generates a dynamic data-ingestion pipeline that is compatible with a Hadoop file system based on the predefined configuration instructions. This dynamic data-ingestion pipeline includes a modular arrangement of operators from a set of operators that includes: an extraction operator for extracting the data of interest from the source, a converter operator for transforming the data, and a quality-checker operator for checking the transformed data. Moreover, the computer system receives the data from the source. Next, the computer system processes the data using the dynamic data-ingestion pipeline as the data is received without storing the data in memory for the purpose of subsequent ingestion processing.


