Data Lineage Service for Heterogeneous Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity of managing large amounts of data across heterogeneous systems in big data projects makes it difficult to understand, manage, and govern data lineage, especially in compliance with regulations, due to the lack of data control and the need for impact and lineage analysis across connected systems.
Innovation Solution
A data lineage service that provides data lineage and impact analysis across a network of heterogeneous systems by creating table-based representations of datasets, transforming lineage information using a unified metadata model, and translating computation specifications into data structures, offering dataset and attribute-level lineage information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in multiple heterogeneous systems for big data projects, then data storage capacity and processing capability are improved, but data lineage management complexity increases
Solution Approach 1:
The patent introduces a data lineage service as an intermediary component that operates between heterogeneous data systems. This service extracts lineage information from various data sources, transforms it into a unified format, and stores it in a centralized lineage database. The intermediary approach allows multiple heterogeneous systems to coexist while providing a standardized interface for lineage management, thereby reducing overall system complexity.
Solution Approach 2:
The patent segments the data lineage management function into distinct modular components: a data lineage extractor that interfaces with various data sources, a transformer that standardizes lineage information, and a storage component that maintains unified lineage data. This segmentation allows each component to be optimized independently and simplifies the overall system architecture by dividing complex management tasks into manageable units.
2Reliability
If data lineage information is extracted and transformed across heterogeneous systems, then data governance and compliance are improved, but processing time and computational resources increase
Solution Approach 1:
The patent implements preliminary action by extracting and transforming data lineage information in advance, before it is needed for governance or compliance operations. The data lineage service continuously monitors data movements and transformations, pre-processing lineage information and storing it in a standardized format. This allows subsequent governance queries and compliance checks to be performed quickly without requiring real-time processing of raw lineage data from heterogeneous systems.
3Stability of the object's composition
If a unified metadata model is used to transform lineage information, then data lineage management consistency is improved, but system integration complexity increases
Solution Approach 1:
The patent employs a data lineage service as an intermediary layer between heterogeneous data systems and the unified metadata model. This service translates diverse lineage information from various sources into the standardized unified format, and vice versa when querying data. The intermediary approach isolates the complexity of integrating with different systems from the unified model itself, allowing consistency to be maintained without directly complicating the core integration architecture.
Data Source
AI summary
Lineage graphs corresponding to data objects are generated. A data object of the data objects is associated with a source dataset table stored at a data lineage server (DLS). The source dataset table includes data of the data object received from a dataset stored at a data source system (DSS). A lineage graph corresponding to the dataset is determined from the lineage graphs. Based on the lineage graph, one or more data lineage structures are provided. The one or more data lineage structures include data from the dataset and from one or more datasets related to the dataset, and define lineage relationships between the data object and one or more data objects corresponding to the one or more datasets.


