Data Proliferation Graph for Lineage Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to effectively model, visualize, and analyze large volumes of data lineage due to the sheer volume of data and numerous locations it passes through, making it difficult for users to identify specific data lineage or objects of interest within data proliferation graphs.
Innovation Solution
A method and system for generating a data proliferation graph that allows users to quickly identify objects of interest by selecting a target data store, identifying downstream or upstream data stores based on metadata tags, and visualizing data lineage through a proliferation graph with adjustable paths and ranking criteria.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data lineage tracking is performed across large volumes of data through numerous locations, then complete data lineage information is obtained, but visualization and analysis become difficult due to the sheer volume of data
Solution Approach 1:
The patent segments the data proliferation graph into multiple columns representing different proliferation levels (degrees of separation from target data store). Each column contains data stores at a specific level, allowing users to navigate and analyze large volumes of data lineage information in manageable segments rather than viewing all data at once.
Solution Approach 2:
The patent transforms the complex data lineage information into a two-dimensional graphical interface where one dimension represents proliferation levels (columns) and the other represents individual data stores (rows within columns). This dimensional organization makes the visualization manageable and analyzable despite the large volume of underlying data.
2Loss of information
If all data stores in the proliferation path are displayed, then complete data lineage is visualized, but users cannot quickly identify specific objects of interest
Solution Approach 1:
The patent applies local quality by allowing users to select and focus on specific sections of the graph (local areas of interest) while maintaining the ability to see the complete data lineage. Users can navigate to specific proliferation levels or data stores of interest without losing the context of the overall data lineage structure.
Solution Approach 2:
The patent introduces an intermediary interface layer between the complete data lineage data and the user view. This interface allows users to select target data stores and navigate through proliferation levels, acting as a mediator that enables focused analysis of specific objects while maintaining access to the complete lineage information.
3Reliability
If detailed data lineage information is collected for security risk mitigation, then risk identification capability is improved, but the complexity of analyzing and navigating the data increases
Solution Approach 1:
The patent implements dynamic navigation capabilities where users can interactively explore the data proliferation graph by selecting target data stores and adjusting the view to focus on specific proliferation levels or paths. This dynamic interaction enables security analysts to efficiently navigate detailed lineage information without being overwhelmed by the complete dataset.
Solution Approach 2:
The patent segments the detailed data lineage information into organized columns by proliferation level, allowing security analysts to systematically review data flow through different stages. This segmentation makes the detailed information more manageable and easier to analyze for security risks while maintaining complete lineage tracking.
Data Source
AI summary
An apparatus, computer-readable medium, and computer-implemented method for generating a data proliferation graph, including receiving a selection of a target data store, identifying a plurality of data stores which have either received data that was previously on the target data store or which have sent data that was subsequently on the target data store, the plurality of data stores being divided into a plurality of proliferation levels corresponding to degrees of separation from the target data store and direction of data propagation relative to the target data store, generating a data proliferation graph, and transmitting at least one portion of the data proliferation graph.


