Inferring Design-Time Lineage from Runtime Artifacts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In distributed computing clusters employing a 'schema on read' approach, such as Apache Hadoop, users lack visibility into the data processing workflow, making it difficult to manage, optimize, and reuse data processing systems effectively due to the absence of predefined schema and hidden design-time information.

Innovation Solution

A technique is introduced to automatically collect, visualize, and utilize upstream and downstream data lineage by inferring design-time information from run-time artifacts, providing users with insights into data processing workflows and enabling the recreation, optimization, and management of data processing systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a 'schema on read' approach is used to allow flexible data processing without predefined schema, then adaptability and ease of data utilization are improved, but visibility into data processing workflows and understanding of data lineage deteriorate

Engineering Contradiction:
Improveflexibility in data processingVSAvoidvisibility into data processing workflow
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent implements feedback by automatically collecting run-time execution information and using it to generate design-time artifacts. The system monitors data processing workflows at runtime, captures metadata about data transformations and lineage, and feeds this information back to create visualizations and documentation that improve workflow visibility and understanding.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary mechanism that bridges the gap between runtime execution and design-time understanding. This intermediary automatically collects execution metadata, processes it into meaningful lineage information, and presents it through visualizations that help users understand data flows without requiring predefined schemas.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If no predefined schema is imposed to enable flexible data loading, then ease of data ingestion is improved, but ability to manage and optimize data processing systems deteriorates

Engineering Contradiction:
Improveease of data loadingVSAvoidability to manage and optimize systems
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling the system to automatically generate design-time artifacts without manual intervention. The system autonomously collects runtime metadata, infers data lineage, and creates visualizations and documentation, eliminating the need for manual schema definition while maintaining system manageability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameter of when schema information is applied - moving from predefined schema-at-write-time to inferred schema-at-read-time. This allows flexible data loading while automatically deriving management information from actual runtime behavior, enabling both ease of ingestion and system optimization.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If design-time information is not collected upfront to enable quick data loading, then productivity is improved, but ability to verify system reliability and optimize performance deteriorates

Engineering Contradiction:
Improvespeed of data processingVSAvoidsystem reliability verification
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by collecting and analyzing runtime execution information to prepare design-time artifacts in advance. The system proactively gathers metadata during execution and uses this information to create lineage visualizations and reliability assessments before they are needed, enabling both fast processing and verification.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical processes of schema definition and reliability verification with automated computational processes. The system uses computational analysis of runtime artifacts to automatically infer data lineage, assess reliability, and generate optimizations, replacing manual planning and design activities.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11663033B2Design-time information based on run-time artifacts in a distributed computing cluster
Publication Date: 2023.05.30 CLOUDERA INC
  • US11663033B2 patent drawing
  • US11663033B2 patent drawing
  • US11663033B2 patent drawing

AI summary

Techniques are disclosed for inferring design-time information based on run-time artifacts generated by services operating in a distributed computing cluster. In an embodiment, a metadata system extracts metadata including run-time artifacts generated by services in a distributed computing cluster while processing a workflow including multiple jobs. The extracted metadata is processed to identify entities and entity relationships which can then be used to generate lineage information. Using the lineage information, the metadata system can infer design-time information associated with the workflow. The inferred design-time information can then be utilized to, for example, recreate the workflow, recreate previous versions of the workflow, optimize the workflow, etc.