Flexible ETL Process for Automated Data Lineage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current ETL processes in data analytics require significant custom modeling and coding, leading to inefficiencies and increased complexity in managing Big Data stores, especially in enterprise environments, where custom development and manual curation of data lineage are necessary, making it difficult to scale and maintain data integration across various sources and systems.

Innovation Solution

The implementation of flexible ETL processes that configure data lineage within a data catalog, allowing for automated data mapping and movement between diverse data sources and target systems, reducing the need for custom coding and enabling scalable data integration without manual curation, thus improving data movement and optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional ETL processes are used with custom modeling and coding, then data integration can be achieved, but system complexity and maintenance effort increase significantly

Engineering Contradiction:
Improveease of data integrationVSAvoidsystem complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent implements a universal ETL framework that can handle multiple data sources, formats, and target systems through a single standardized platform. The system provides pre-built connectors, templates, and components that work across diverse scenarios, eliminating the need for custom modeling and coding for each integration project while reducing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If custom development is performed for data integration, then specific intelligence needs are met, but scalability and maintainability are reduced

Engineering Contradiction:
Improveadaptability to intelligence needsVSAvoidscalability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements a dynamic ETL system with configurable parameters, templates, and components that can be adjusted to meet specific intelligence needs without requiring custom development. The system allows users to dynamically select and configure appropriate data sources, transformations, and targets through a standardized interface, enabling both adaptability and scalability simultaneously.

Inventive Principle:
Principle #15Dynamics

3Reliability

If manual curation of data lineage is performed, then data tracking accuracy is maintained, but time and resource consumption increase

Engineering Contradiction:
Improvedata lineage accuracyVSAvoidtime for lineage curation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements an automated data lineage tracking system that self-documented data flows, transformations, and dependencies throughout the ETL process. The system automatically captures metadata, tracks data provenance, and generates lineage reports without requiring manual curation, thereby maintaining accuracy while eliminating time and resource consumption associated with manual tracking.

Inventive Principle:
Principle #25Self-service

4Ease of operation

If ETL processes are simplified to reduce complexity, then ease of operation improves, but measurement precision and data transformation accuracy may be compromised

Engineering Contradiction:
Improveease of ETL operationVSAvoiddata transformation accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent introduces standardized intermediaries in the form of pre-built transformation components, templates, and validation rules that bridge the gap between simplified operation and precise data transformation. These intermediaries encapsulate complex transformation logic while presenting simple interfaces to users, ensuring both ease of operation and manufacturing precision simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240386027A1Flexible extract, transform, and load (ETL) process
Publication Date: 2024.11.21 HIGHCHEM SRO
  • US20240386027A1 patent drawing
  • US20240386027A1 patent drawing
  • US20240386027A1 patent drawing

AI summary

Disclosed herein are systems, methods, computing devices, and computer-readable media. For example, in some embodiments, a method for extract, transform and load (ETL) processing executed by a processing device, the method comprising: receiving, from a data catalog, field mapping between application data and a target schema; receiving from an ETL processing queue, signal comprising metadata and indicating that a record or a data file related to the application data is ready for processing; determining source data by processing the metadata to identify a location of the record or the data file and retrieving the record or the data file from the identified location; and providing source data to a table defined according to the target schema.