Filtering Data Lineage Diagrams Using Dimensional Tags
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data lineage diagrams in data processing systems become complex due to numerous nodes and lineage paths, making it difficult for users to focus on specific, relevant information, especially when dealing with normalization and de-normalization operations, as existing filtering methods limit the view to direct lineage or impact paths, failing to selectively exclude unwanted nodes.
Innovation Solution
The implementation of lineage tags with user-defined dimensions allows for filtering data lineage diagrams, enabling selective inclusion or exclusion of nodes based on tag associations, prioritizing lower-level node tags over higher-level ones, and traversing lineage paths to generate a filtered representation that targets specific user-defined dimensions without limiting the view to direct lineage or impact paths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data lineage diagrams include all nodes and lineage paths to ensure completeness, then the diagram provides comprehensive information, but the complexity increases making it difficult for users to focus on specific relevant information
Solution Approach 1:
The patent applies segmentation by dividing the data lineage diagram into multiple filtered views based on user-selected dimensions and criteria. Instead of presenting all nodes and paths simultaneously, the system segments the lineage information into manageable subsets that can be selectively displayed, reducing visual complexity while maintaining access to complete information through multiple filtered perspectives
Solution Approach 2:
The patent introduces filtering dimensions as an additional layer of organization beyond the basic lineage structure. By adding dimensional filters (such as data type, processing stage, or business function), users can navigate the lineage diagram through multiple dimensions, effectively reducing complexity by viewing the same comprehensive data through different organizational lenses
2Device complexity
If filtering methods limit the view to direct lineage or impact paths, then the diagram becomes simpler to view, but the ability to selectively exclude unwanted nodes is limited
Solution Approach 1:
The patent implements dynamic filtering where users can selectively apply different filtering criteria to different dimensions of the lineage diagram. The filtering mechanism is not static but adapts to user preferences, allowing dynamic adjustment of which nodes to include or exclude based on multiple configurable parameters simultaneously, providing both simplicity and selectivity
3Manufacturing precision
If normalization and de-normalization operations are included in the lineage diagram, then the diagram accurately represents data transformations, but the number of nodes increases adding to the complexity
Solution Approach 1:
The patent applies local quality by allowing users to apply different filtering treatments to different regions or types of nodes within the lineage diagram. Specifically, normalization and de-normalization nodes can be selectively filtered based on user preferences, enabling accurate representation of data transformations where needed while reducing complexity in other areas through selective node exclusion
Data Source
AI summary
Managing lineage information includes processing a specification of a directed graph to associate nodes with information for processing requests for a representation of data lineage. The processing includes: identifying a first set of one or more nodes of the directed graph corresponding to normalizing data elements being stored in a data store and de-normalizing data elements being retrieved from the data store; and associating a first plurality of nodes connected to the first set of one or more nodes and a second plurality of nodes connected to the first set of one or more nodes with at least one tag identifier having a plurality of possible tag values, where the number of possible tag values is at least as large as the number of data elements being normalized, and where nodes representing different data elements in a de-normalized record are associated with different values of the tag identifier.


