Data Proliferation Graph for Lineage Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods struggle to effectively model, visualize, and analyze large volumes of data lineage due to the sheer volume of data and numerous locations it passes through, making it difficult for users to identify specific data lineage or objects of interest within data proliferation graphs.

Innovation Solution

A method and system for generating a data proliferation graph that allows users to quickly identify objects of interest by selecting a target data store, identifying downstream or upstream data stores based on metadata tags, and visualizing data lineage through a proliferation graph with adjustable paths and ranking criteria.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data lineage tracking is performed across large volumes of data through numerous locations, then complete data lineage information is obtained, but visualization and analysis become difficult due to the sheer volume of data

Engineering Contradiction:
Improvedata lineage informationVSAvoidvisualization complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the data proliferation graph into multiple columns representing different proliferation levels (degrees of separation from target data store). Each column contains data stores at a specific level, allowing users to navigate and analyze large volumes of data lineage information in manageable segments rather than viewing all data at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the complex data lineage information into a two-dimensional graphical interface where one dimension represents proliferation levels (columns) and the other represents individual data stores (rows within columns). This dimensional organization makes the visualization manageable and analyzable despite the large volume of underlying data.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If all data stores in the proliferation path are displayed, then complete data lineage is visualized, but users cannot quickly identify specific objects of interest

Engineering Contradiction:
Improvedata lineage completenessVSAvoidobject identification difficulty
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies local quality by allowing users to select and focus on specific sections of the graph (local areas of interest) while maintaining the ability to see the complete data lineage. Users can navigate to specific proliferation levels or data stores of interest without losing the context of the overall data lineage structure.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces an intermediary interface layer between the complete data lineage data and the user view. This interface allows users to select target data stores and navigate through proliferation levels, acting as a mediator that enables focused analysis of specific objects while maintaining access to the complete lineage information.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If detailed data lineage information is collected for security risk mitigation, then risk identification capability is improved, but the complexity of analyzing and navigating the data increases

Engineering Contradiction:
Improvesecurity risk mitigationVSAvoiddata navigation ease
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements dynamic navigation capabilities where users can interactively explore the data proliferation graph by selecting target data stores and adjusting the view to focus on specific proliferation levels or paths. This dynamic interaction enables security analysts to efficiently navigate detailed lineage information without being overwhelmed by the complete dataset.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent segments the detailed data lineage information into organized columns by proliferation level, allowing security analysts to systematically review data flow through different stages. This segmentation makes the detailed information more manageable and easier to analyze for security risks while maintaining complete lineage tracking.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11134096B2Method, apparatus, and computer-readable medium for generating data proliferation graph
Publication Date: 2021.09.28 INFORMATICA CORP
  • US11134096B2 patent drawing
  • US11134096B2 patent drawing
  • US11134096B2 patent drawing

AI summary

An apparatus, computer-readable medium, and computer-implemented method for generating a data proliferation graph, including receiving a selection of a target data store, identifying a plurality of data stores which have either received data that was previously on the target data store or which have sent data that was subsequently on the target data store, the plurality of data stores being divided into a plurality of proliferation levels corresponding to degrees of separation from the target data store and direction of data propagation relative to the target data store, generating a data proliferation graph, and transmitting at least one portion of the data proliferation graph.