Lineage Engine Data Flow Direction Determination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face challenges in effectively utilizing vast amounts of data due to its sheer quantity and diversity, making it difficult to identify relationships between disparate data sources, which can lead to errors in manual determination and inefficient visualization.
Innovation Solution
A data model is created using a lineage engine that represents data types and relationships as nodes and edges, with schema content providing hierarchy indicators to automatically determine data flow directions and generate shortcut edges, enabling unambiguous traversal and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual determination of relationships between data types is used, then flexibility in handling varied data types is maintained, but error rate increases and time consumption increases
Solution Approach 1:
The system automatically determines relationships between data types by analyzing schema content and generating data flow directions without requiring manual intervention. The lineage engine self-processes the schema to identify parent-child relationships, data flow paths, and hierarchy indicators, eliminating the need for manual analysis while improving accuracy and reducing time consumption.
Solution Approach 2:
The patent replaces manual mechanical analysis of data relationships with an automated computational system. The lineage engine uses algorithmic processing to parse schema content, generate data models, and determine relationships automatically, substituting human manual determination with automated mechanical processing that is both faster and more accurate.
2Productivity
If automated data flow direction determination is implemented, then time for relationship determination is reduced, but system complexity increases
Solution Approach 1:
The system segments the data modeling process into distinct manageable components: schema content parsing, data type identification, relationship determination, and data flow direction generation. The lineage engine processes schema content in discrete steps to generate data models with defined relationships, making the complex automated system modular and maintainable.
Solution Approach 2:
The patent introduces schema content as an intermediary representation that bridges raw data structures and the final data model. This intermediate layer standardizes the input format and provides structured information that the lineage engine can process systematically, reducing the complexity of directly processing diverse data formats while maintaining automation capabilities.
3Measurement precision
If hierarchy indicators are added to schema content, then data flow direction accuracy is improved, but schema content complexity increases
Solution Approach 1:
The patent applies hierarchy indicators locally at specific points in the data model where data flow direction needs to be explicitly defined. Rather than adding complexity throughout the entire schema, the indicators are placed only where necessary to disambiguate data flow paths, providing precise directional information only at critical decision points in the data hierarchy.
Data Source
AI summary
Embodiments are directed to managing data. A data model that includes data type nodes and relationship edges may be provided. Other data types and other data type relationships may be provided and included in the data model. If a portion of the nodes in the data model may be downstream of leaf nodes in the graphlet: the data model may be traversed to visit the downstream nodes; shortcut edges may be generated to each downstream node associated with shortcut nodes. If a second portion of the nodes in the data model may be upstream of the leaf nodes: the data model may be traversed upwards from the leaf nodes; other shortcut edges may be generated to each node visited in the upwards traversal associated with shortcut nodes.


