Data Lineage Scoring and Filtering Mechanism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data managers and end-users face inefficiencies in navigating comprehensive data lineage information, as existing tools provide detailed and complex data lineage analysis that is not always necessary for their purposes.
Innovation Solution
A method is introduced that assigns scores to data assets based on predefined criteria and filters results to present only relevant data lineage information, using a scoring function and filtering criteria to determine inclusion in the output, thereby simplifying the representation of data lineage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If comprehensive data lineage analysis is performed to provide complete information about data sources, transformations, and destinations, then the completeness and reliability of data lineage information is improved, but the complexity and volume of information presented to users increases significantly
Solution Approach 1:
The patent applies parameter changes by introducing scoring criteria that evaluate data assets based on multiple attributes (business relevance, data quality, usage frequency, etc.). This transforms the raw comprehensive data lineage information into a scored and prioritized representation, allowing users to focus on high-value assets while maintaining access to complete information when needed.
Solution Approach 2:
The patent implements local quality by applying different scoring weights and criteria to different types of data assets based on their specific characteristics. Each data asset is evaluated according to its own relevance to business objectives, data quality metrics, and usage patterns, rather than applying a uniform filtering approach to all assets.
2Reliability
If comprehensive data lineage analysis is performed to enumerate all databases, files, systems, and transformations, then the completeness of data lineage information is improved, but the time and effort required for users to sift through the information increases
Solution Approach 1:
The patent applies preliminary action by pre-calculating scores for all data assets in the lineage path before presenting information to users. The scoring function evaluates business relevance, data quality, and other criteria in advance, so that when users query data lineage, they immediately receive prioritized results rather than having to filter through comprehensive information manually.
Solution Approach 2:
The patent transforms the presentation of data lineage information by introducing scoring parameters that rank assets according to their importance. This parameter change allows the system to maintain complete lineage information internally while presenting a time-efficient, prioritized view to users based on configurable scoring criteria.
3Reliability
If all data assets along data lineage paths are presented to users, then the completeness of information is improved, but the usefulness and relevance of the information to specific user purposes decreases
Solution Approach 1:
The patent implements local quality by tailoring the presentation of data lineage information to specific user needs and contexts. Different users or user groups can configure scoring criteria that reflect their particular objectives, such as focusing on data quality for compliance users or on usage frequency for operational users, making the information more useful for their specific purposes.
Solution Approach 2:
The patent applies dynamics by making the data lineage presentation adaptive and configurable. Users can modify scoring criteria and weights based on their evolving needs, and the system dynamically adjusts which assets are highlighted or filtered in the results, rather than providing a static comprehensive view that may not align with current user objectives.
Data Source
AI summary
Presenting data lineage information by assigning a score to a data asset along a path between a data source and a data destination, where a predefined scoring function is applied to a characteristic of the data asset, and presenting via a computer-controlled output medium a description of the data source, the data destination, and the path between the data source and the data destination, where the description includes the data asset if the score meets predefined inclusion criteria.


