Data Lineage Tracking via Automated Dataset Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Enterprises face challenges in efficiently organizing and reporting vast amounts of data, particularly in complying with regulatory requirements like FR 2052a, due to manual intensive efforts and lack of efficient data pipeline tracking.

Innovation Solution

A method and system for creating dataflows that provide a 360-degree view of data by classifying source datasets based on use cases, applying data processing rules, and generating a curated dataset with associated data lineage, which is then visualized and stored in a network-accessible data store.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual research and analysis of dataflows is performed by technology and business groups, then data lineage can be tracked, but the process becomes manually intensive and resource-consuming

Engineering Contradiction:
Improvedata lineage tracking accuracyVSAvoiddata organization efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables automated self-service data lineage tracking by implementing metadata extraction and classification mechanisms that automatically monitor and document dataflows without requiring manual intervention from technology or business groups, thereby maintaining reliability while significantly improving productivity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of researching and analyzing dataflows with an automated computer-implemented system that uses metadata extraction, classification, and visualization technologies to automatically track and document data lineage, eliminating manual efforts while preserving tracking accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If data is compiled from multiple different sources to comply with reporting requirements, then comprehensive reporting can be achieved, but the effort becomes manually intensive

Engineering Contradiction:
Improvereporting accuracyVSAvoiddata compilation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary classification and organization of metadata from multiple data sources before actual reporting is needed, creating a structured foundation that enables rapid and accurate report generation without manual compilation efforts when reporting requirements arise

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a universal metadata classification framework that can handle diverse data sources and multiple reporting requirements simultaneously, allowing the same automated system to serve various reporting needs (FR 2052a, general liquidity reporting, analytics, debt reporting) without requiring separate manual processes for each

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12326854B1Data landing zone framework
Publication Date: 2025.06.10 WELLS FARGO BANK NA
  • US12326854B1 patent drawing
  • US12326854B1 patent drawing
  • US12326854B1 patent drawing

AI summary

Systems and methods for providing a data lineage are described. According to some examples, a method may include receiving a source dataset corresponding to a use case scenario. The method may further include classifying the source dataset based on the use case scenario to create a classified dataset and then applying a set of data processing rules to the classified dataset to create a curated dataset. The method may further include receiving a reporting data object indicating a relationship between the source dataset and the curated dataset. The relationship can describe a data lineage of the source dataset. The method may further include outputting display data associated with the reporting data object.