Metadata-Driven Data Ingestion Framework for Real-Time Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data ingestion practices in distributed data frameworks, such as Hadoop, are cumbersome and time-consuming due to the need for multiple tool configurations and high latency, which hinders real-time data availability.

Innovation Solution

A system for data ingestion in a distributed processing framework that includes a processor executing instructions to access data from sources, identify metadata attributes, calculate derived attributes, and generate enriched data, creating a physical view for efficient data processing and real-time availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple specialized tools are used to handle different data flows, then data processing coverage is improved, but system complexity and configuration difficulty increase

Engineering Contradiction:
Improvedata processing coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal data ingestion framework that can handle multiple types of data flows (batch, real-time, streaming) through a single unified system. The framework uses configurable processors that can be dynamically adjusted to handle different data types and processing requirements, eliminating the need for multiple specialized tools while maintaining comprehensive data processing coverage.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The unified data ingestion framework is segmented into modular, configurable processors that can be independently configured and combined. Each processor handles specific data processing tasks but can be dynamically configured through metadata to handle different data types, allowing the system to maintain versatility while reducing overall complexity through standardized modular components.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If multiple specialized tools with specific coding configurations are implemented, then data processing capability is improved, but time-to-market and implementation time increase

Engineering Contradiction:
Improvedata processing capabilityVSAvoidtime-to-market
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The framework employs pre-configured processor templates and metadata schemas that are prepared in advance. When new data processing requirements arise, the system can quickly instantiate and configure processors using these pre-prepared templates rather than building from scratch, significantly reducing implementation time while maintaining processing capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The data ingestion framework enables self-service configuration through metadata-driven processing. Users can define data processing requirements using standardized metadata schemas, and the system automatically configures the appropriate processors without requiring extensive manual coding or specialized configuration expertise, thereby accelerating time-to-market.

Inventive Principle:
Principle #25Self-service

3Reliability

If existing data ingestion tools are used, then data processing is achieved, but latency is high and real-time data availability is limited

Engineering Contradiction:
Improvedata processing reliabilityVSAvoiddata ingestion latency
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The framework implements dynamic processor configuration that allows processing behavior to be adjusted in real-time based on data characteristics and system conditions. Processors can dynamically switch between batch and streaming modes, adjust processing intensity, and optimize latency based on incoming data requirements, enabling low-latency real-time processing while maintaining reliability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters dynamically based on metadata definitions and data flow characteristics. By adjusting parameters such as processing mode (batch vs. streaming), buffer sizes, and processing intensity based on the specific data requirements defined in metadata, the framework achieves low-latency real-time processing for appropriate data types while maintaining reliable processing for all data types.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11481451B2Metadata driven combined real-time and batch data ingestion framework with real-time multi-view generation
Publication Date: 2022.10.25 JPMORGAN CHASE BANK NA
  • US11481451B2 patent drawing
  • US11481451B2 patent drawing
  • US11481451B2 patent drawing

AI summary

Systems and methods for ingesting and enhancing data in a distributed processing framework. The system includes at least a data ingestion system configured to access data or datasets from one or more data sources. The data is accessed via the data ingestion system and includes metadata defining a plurality of attributes. The attributes are identified in the metadata, via the data ingestion system, and may be applied to the data or dataset for enhancing the data or dataset. Application of the attributes to the data results in enhancements that may include joining the data, enriching the data, or other enhancements accomplished via manipulation of the data via the data ingestion system.