Data Lineage Tracking via Automated Dataset Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprises face challenges in efficiently organizing and reporting vast amounts of data, particularly in complying with regulatory requirements like FR 2052a, due to manual intensive efforts and lack of efficient data pipeline tracking.
Innovation Solution
A method and system for creating dataflows that provide a 360-degree view of data by classifying source datasets based on use cases, applying data processing rules, and generating a curated dataset with associated data lineage, which is then visualized and stored in a network-accessible data store.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual research and analysis of dataflows is performed by technology and business groups, then data lineage can be tracked, but the process becomes manually intensive and resource-consuming
Solution Approach 1:
The system enables automated self-service data lineage tracking by implementing metadata extraction and classification mechanisms that automatically monitor and document dataflows without requiring manual intervention from technology or business groups, thereby maintaining reliability while significantly improving productivity
Solution Approach 2:
The patent replaces the mechanical manual process of researching and analyzing dataflows with an automated computer-implemented system that uses metadata extraction, classification, and visualization technologies to automatically track and document data lineage, eliminating manual efforts while preserving tracking accuracy
2Reliability
If data is compiled from multiple different sources to comply with reporting requirements, then comprehensive reporting can be achieved, but the effort becomes manually intensive
Solution Approach 1:
The system performs preliminary classification and organization of metadata from multiple data sources before actual reporting is needed, creating a structured foundation that enables rapid and accurate report generation without manual compilation efforts when reporting requirements arise
Solution Approach 2:
The patent implements a universal metadata classification framework that can handle diverse data sources and multiple reporting requirements simultaneously, allowing the same automated system to serve various reporting needs (FR 2052a, general liquidity reporting, analytics, debt reporting) without requiring separate manual processes for each
Data Source
AI summary
Systems and methods for providing a data lineage are described. According to some examples, a method may include receiving a source dataset corresponding to a use case scenario. The method may further include classifying the source dataset based on the use case scenario to create a classified dataset and then applying a set of data processing rules to the classified dataset to create a curated dataset. The method may further include receiving a reporting data object indicating a relationship between the source dataset and the curated dataset. The relationship can describe a data lineage of the source dataset. The method may further include outputting display data associated with the reporting data object.


