Metadata Visualization for Data Lineage Tracing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In complex data processing systems, tracing the lineage of data elements and understanding their relationships is challenging due to the complexity of data sources and transformations, making it difficult for users to track the origins and impacts of datasets.
Innovation Solution
A method is implemented to store metadata in a data storage system, compute summary data characterizing metadata objects, and generate visual representations of diagrams showing nodes representing metadata objects and their relationships, with superimposed characteristics such as quality, update status, and source information, allowing users to visualize and interact with this information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data are pulled from many different data sources into a central repository with transformations, then data processing capability is improved, but tracing data lineage and understanding relationships becomes difficult
Solution Approach 1:
The patent introduces metadata as an intermediary layer between raw data and user queries. Metadata objects store information about data sources, transformations, and relationships, allowing users to trace data lineage without directly analyzing complex data processing systems. The visual representation system acts as another intermediary, translating complex metadata relationships into intuitive graphical displays.
Solution Approach 2:
The patent replaces manual tracing of data lineage through complex data processing systems with an automated visual representation system. The system automatically generates graphical diagrams showing data relationships, sources, and transformations, eliminating the need for users to manually investigate complex data processing paths.
2Loss of information
If detailed metadata about data sources and transformations is stored, then data lineage information is improved, but visualization complexity and information overload increase
Solution Approach 1:
The patent segments metadata into distinct metadata objects, each representing specific aspects of data lineage (sources, transformations, relationships). The visual representation system further segments information by displaying only relevant characteristics near corresponding nodes in the diagram, rather than showing all metadata at once. This segmentation reduces visualization complexity while preserving complete lineage information.
Solution Approach 2:
The patent applies local quality by displaying different characteristics of metadata objects at different locations in the visual representation. Summary data characterizing specific aspects (e.g., data quality, source information, transformation details) are displayed in proximity to the relevant nodes or edges in the diagram, allowing users to access detailed information locally rather than presenting a uniform complex view everywhere.
3Ease of operation
If summary data characterizing metadata objects is computed and stored, then data understanding is improved, but processing time and computational resources increase
Solution Approach 1:
The patent computes and stores summary data characterizing metadata objects in advance, before users need to query or visualize the data. This preliminary computation of summary information (such as data quality metrics, source summaries, or transformation descriptions) reduces the time required to generate visual representations and answer user queries, as the heavy computational work has already been performed.
Data Source
AI summary
In general, metadata is stored in a data storage system. Summary data identifying one or more characteristics of each of multiple metadata objects stored in the data storage system is computed, and the summary data characterizing a given metadata object in association with the given metadata object is stored. A visual representation is generated of a diagram including nodes representing respective metadata objects and relationships among the nodes. Generating the visual representation includes superimposing a representation of a characteristic identified by the summary data characterizing a given metadata object in proximity to the node representing the given metadata object.


