Dynamic Data-Ingestion Pipeline for Hadoop

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Handling high volumes of heterogeneous data from various sources poses challenges in computer systems, including increased complexity, high costs, and degraded performance due to the need for significant memory and offline processing.

Innovation Solution

A dynamic data-ingestion pipeline is generated using a modular arrangement of operators, allowing real-time processing without storing data in memory, and adaptable to various data sources, including Hadoop file systems, to ingest and transform data efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If significant memory is used to store large amounts of data from different data sources, then data storage capacity is improved, but system cost and complexity increase

Engineering Contradiction:
Improvedata storage capacityVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary data fields from incoming data streams rather than storing entire records. The schema-based extraction mechanism identifies and pulls only relevant fields, significantly reducing memory requirements while maintaining data utility for processing

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The data processing system is segmented into multiple independent components: data ingestion services, schema validation modules, extraction operators, and output writers. Each component handles specific tasks independently, reducing overall system complexity while maintaining high data storage capacity

Inventive Principle:
Principle #1Segmentation

2Loss of energy

If offline or non-real-time processing is used, then system cost is reduced, but performance and user experience degrade

Engineering Contradiction:
Improvesystem costVSAvoiddata processing performance
Core Design Contradiction:
Loss of energyVSProductivity

Solution Approach 1:

The patent implements dynamic data extraction where the system adapts its processing behavior based on incoming data characteristics and schema definitions. The extraction operators dynamically adjust which fields to pull based on schema validation results, enabling real-time processing without static rigid configurations

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

A schema validation intermediary layer is introduced between data ingestion and extraction. This intermediary validates incoming data against defined schemas and guides the extraction process, enabling real-time processing decisions without requiring complete offline analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If separate data processing is implemented for each data source, then data source compatibility is improved, but system complexity increases

Engineering Contradiction:
Improvedata source compatibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal data extraction framework that handles multiple data sources through a common interface. The schema-based extraction operators work uniformly across different data sources (Hadoop, relational databases, NoSQL), providing multi-functionality without requiring separate processing logic for each source

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses schema copying and template matching to handle different data sources. Instead of creating custom processing logic for each source, the system copies and adapts schema templates to match various data source structures, enabling compatibility through pattern replication rather than custom implementation

Inventive Principle:
Principle #26Copying

4Speed

If high velocity data integration is implemented, then data processing speed is improved, but system cost increases

Engineering Contradiction:
Improvedata integration velocityVSAvoidsystem cost
Core Design Contradiction:
SpeedVSLoss of energy

Solution Approach 1:

The patent applies partial action by extracting only the necessary subset of fields from incoming data rather than processing and storing all data. This selective extraction achieves high velocity integration for the critical data fields without the cost of processing entire data records

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10122783B2Dynamic data-ingestion pipeline
Publication Date: 2018.11.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10122783B2 patent drawing
  • US10122783B2 patent drawing
  • US10122783B2 patent drawing

AI summary

In order to ingest data from an arbitrary source in a set of sources, a computer system accesses predefined configuration instructions. Then, the computer system generates a dynamic data-ingestion pipeline that is compatible with a Hadoop file system based on the predefined configuration instructions. This dynamic data-ingestion pipeline includes a modular arrangement of operators from a set of operators that includes: an extraction operator for extracting the data of interest from the source, a converter operator for transforming the data, and a quality-checker operator for checking the transformed data. Moreover, the computer system receives the data from the source. Next, the computer system processes the data using the dynamic data-ingestion pipeline as the data is received without storing the data in memory for the purpose of subsequent ingestion processing.