Data Lineage Tracking via Trace Log Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data lineage tracking in heterogeneous computing environments is challenging due to the heterogeneity of systems and complexity, leading to obscured, gap-filled, or lost data lineage information, which affects the confidence of decisions made in enterprise environments.

Innovation Solution

A data lineage tracking system and method that captures and manages metadata, tracks and produces data flows, infers reasoning for changes, and reports data lineage information across heterogeneous platforms and applications using trace logs, clustering similar log entries, and analyzing mappings to determine data lineage flow.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data lineage tracking is implemented in heterogeneous computing environments, then data provenance information can be captured, but the heterogeneity of systems and complexity cause the lineage information to become obscured, gap-filled, or lost

Engineering Contradiction:
Improvedata lineage informationVSAvoidsystem heterogeneity and complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces intermediary components including a data lineage tracker that intercepts and records data flow between heterogeneous systems, a standardized metadata schema that mediates between different data formats, and a lineage repository that serves as a central intermediary for storing and retrieving lineage information. These intermediaries translate and normalize data exchanges across diverse systems, preventing information loss despite system heterogeneity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements universal data lineage tracking mechanisms that can operate across multiple heterogeneous computing systems simultaneously. The lineage tracker and metadata schema are designed to be system-agnostic, handling various data types, formats, and processing operations through a unified framework that adapts to different environments while maintaining consistent lineage capture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If manual tracking of data lineage is performed, then some data provenance information can be recorded, but the process is time-consuming and resource-intensive

Engineering Contradiction:
Improvedata lineage informationVSAvoidtime and resources for tracking
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements automated data lineage tracking where the system self-monitors and self-documented data flows without requiring manual intervention. The lineage tracker automatically intercepts data operations, the metadata schema auto-generates lineage records, and the system self-updates the lineage repository, eliminating the need for manual tracking efforts while maintaining comprehensive lineage information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent establishes feedback loops where the automated lineage tracking system continuously monitors data flows, validates lineage information, and adjusts tracking mechanisms based on detected patterns and anomalies. This automated feedback mechanism ensures accurate lineage capture while minimizing manual oversight requirements.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9659042B2Data lineage tracking
Publication Date: 2017.05.23 ACCENTURE GLOBAL SERVICES LTD
  • US9659042B2 patent drawing
  • US9659042B2 patent drawing
  • US9659042B2 patent drawing

AI summary

A data lineage tracking system may include a memory storing a module comprising machine readable instructions to obtain trace log entries representing an interaction with, a manipulation of, and/or a creation of a data value. The data lineage tracking system may further include machine readable instructions to select the trace log entries that are associated with commands performed by an application, cluster similar trace log entries from the selected trace log entries, and analyze mappings between the clustered trace log entries to determine data lineage flow associated with the data value.