Enterprise Data Lineage Tracking via Automated Metadata Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems lack robust tracking capabilities for data lineage across multiple applications and computing systems, leading to inaccurate and outdated data lineage information, which hinders the validation of data transformations and contributes to systematic data tracking deficiencies in enterprise organizations.
Innovation Solution
An enterprise computing system with a data testing module and an asset lineage module that tracks and generates data lineage information by storing reported data lineage, analyzing test case metadata, and creating data lineage maps to accurately depict relationships between source and target data elements, ensuring accurate validation and tracking of data transformations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual entry of data lineage information is used, then ease of operation is improved, but reliability deteriorates due to inaccurate recordation of relationships between source and target data elements
Solution Approach 1:
The system enables self-service automated tracking of data lineage by having applications and data systems automatically report their own data transformation metadata to the centralized platform, eliminating manual entry while ensuring accurate and reliable data lineage information through self-reported verified data
Solution Approach 2:
The patent replaces the mechanical manual entry system with an automated computational system that uses metadata extraction, data flow analysis, and algorithmic tracking to automatically capture and verify data lineage relationships, substituting human-operated mechanical processes with automated electronic systems
2Reliability
If automated tracking of data lineage across multiple applications is implemented, then reliability is improved, but device complexity increases
Solution Approach 1:
The platform provides multi-functional capabilities including automated data lineage tracking, data quality monitoring, metadata management, and validation services within a single unified system, allowing one complex system to perform multiple functions that would otherwise require separate systems
Solution Approach 2:
The patent introduces a centralized data lineage platform as an intermediary layer between source systems and target systems, which mediates the collection, normalization, and verification of data transformation metadata, simplifying the overall architecture by providing a single point of coordination rather than direct complex point-to-point tracking
3Device complexity
If manual data lineage reporting is used, then device complexity is reduced, but loss of information increases due to outdated and invalid data lineage information
Solution Approach 1:
The system implements continuous feedback mechanisms where data lineage information is automatically collected, verified, and updated in real-time as data transformations occur, ensuring that the most current and accurate lineage information is maintained without requiring complex manual intervention
Solution Approach 2:
The patent performs preliminary automated validation and verification of data lineage information at the point of data transformation, checking the accuracy and completeness of metadata before it is recorded, thereby preventing loss of information by ensuring data quality upfront rather than requiring complex post-processing
4Productivity
If automated test case generation from data lineage metadata is implemented, then productivity is improved, but measurement precision requirements increase
Solution Approach 1:
The system performs preliminary automated validation of data quality rules and test criteria before generating test cases, ensuring that the metadata used for test generation meets precision requirements, thereby enabling rapid test case development without sacrificing measurement precision
Solution Approach 2:
The patent replaces manual test case development with automated generation systems that use algorithms to create test cases from validated data lineage metadata, substituting human analysts with computational systems that can rapidly generate precise test cases while maintaining high measurement precision through automated validation rules
Data Source
AI summary
A computing system for managing and mapping source data and target data associated with a data transformation analyzes data quality testing data. Source data and target data include the data elements, data structures, and storage mechanisms for data associated with a data transformation. The computing system analyzes the data quality testing data for validation of the associated data transformation. The computing system identifies source data for input to the data transformation and target data for the result of the data transformation. The computing system stores identifiers associated with the source data and target data and records validated data lineage information for the data transformation. Based on a configuration, the computing system generates a data lineage map indicating the relationships between the source data and the target data associated with a number of data transformations that occur within the computing system.


