Inferring Design-Time Lineage from Runtime Artifacts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed computing clusters employing a 'schema on read' approach, such as Apache Hadoop, users lack visibility into the data processing workflow, making it difficult to manage, optimize, and reuse data processing systems effectively due to the absence of predefined schema and hidden design-time information.
Innovation Solution
A technique is introduced to automatically collect, visualize, and utilize upstream and downstream data lineage by inferring design-time information from run-time artifacts, providing users with insights into data processing workflows and enabling the recreation, optimization, and management of data processing systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a 'schema on read' approach is used to allow flexible data processing without predefined schema, then adaptability and ease of data utilization are improved, but visibility into data processing workflows and understanding of data lineage deteriorate
Solution Approach 1:
The patent implements feedback by automatically collecting run-time execution information and using it to generate design-time artifacts. The system monitors data processing workflows at runtime, captures metadata about data transformations and lineage, and feeds this information back to create visualizations and documentation that improve workflow visibility and understanding.
Solution Approach 2:
The patent introduces an intermediary mechanism that bridges the gap between runtime execution and design-time understanding. This intermediary automatically collects execution metadata, processes it into meaningful lineage information, and presents it through visualizations that help users understand data flows without requiring predefined schemas.
2Ease of manufacture
If no predefined schema is imposed to enable flexible data loading, then ease of data ingestion is improved, but ability to manage and optimize data processing systems deteriorates
Solution Approach 1:
The patent implements self-service by enabling the system to automatically generate design-time artifacts without manual intervention. The system autonomously collects runtime metadata, infers data lineage, and creates visualizations and documentation, eliminating the need for manual schema definition while maintaining system manageability.
Solution Approach 2:
The patent changes the parameter of when schema information is applied - moving from predefined schema-at-write-time to inferred schema-at-read-time. This allows flexible data loading while automatically deriving management information from actual runtime behavior, enabling both ease of ingestion and system optimization.
3Productivity
If design-time information is not collected upfront to enable quick data loading, then productivity is improved, but ability to verify system reliability and optimize performance deteriorates
Solution Approach 1:
The patent applies preliminary action by collecting and analyzing runtime execution information to prepare design-time artifacts in advance. The system proactively gathers metadata during execution and uses this information to create lineage visualizations and reliability assessments before they are needed, enabling both fast processing and verification.
Solution Approach 2:
The patent replaces manual mechanical processes of schema definition and reliability verification with automated computational processes. The system uses computational analysis of runtime artifacts to automatically infer data lineage, assess reliability, and generate optimizations, replacing manual planning and design activities.
Data Source
AI summary
Techniques are disclosed for inferring design-time information based on run-time artifacts generated by services operating in a distributed computing cluster. In an embodiment, a metadata system extracts metadata including run-time artifacts generated by services in a distributed computing cluster while processing a workflow including multiple jobs. The extracted metadata is processed to identify entities and entity relationships which can then be used to generate lineage information. Using the lineage information, the metadata system can infer design-time information associated with the workflow. The inferred design-time information can then be utilized to, for example, recreate the workflow, recreate previous versions of the workflow, optimize the workflow, etc.


