Provenance Storage Prioritization for Distributed Data Pipelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Provenance systems in distributed cloud environments face challenges in managing vast arrays of configuration and policy combinations, leading to cumbersome handling and inefficient storage of provenance information, which is crucial for data processing pipelines.

Innovation Solution

A method and system that derive declarative intents representing configuration requirements and priority levels for storing provenance information, estimate storage capacity, and reduce data amount by compressing, removing indexing, or moving lower priority data to remote storage when thresholds are met, ensuring efficient resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If provenance systems capture comprehensive provenance information for data processing pipelines, then the completeness and usefulness of provenance data is improved, but the storage capacity consumption increases significantly

Engineering Contradiction:
Improvecompleteness of provenance informationVSAvoidstorage capacity consumption
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies local quality by assigning different retention policies to different provenance entities based on their priority levels. High-priority entities maintain complete provenance information locally, while low-priority entities have their information moved to remote storage or compressed, thereby optimizing storage usage while preserving information quality where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments provenance information management into priority-based categories. By dividing provenance entities into high-priority and low-priority groups, the system can apply different storage strategies to each segment, reducing overall storage consumption while maintaining essential information.

Inventive Principle:
Principle #1Segmentation

2Reliability

If the system stores detailed provenance information for all data processing pipelines, then the reliability of provenance data is improved, but the device complexity increases due to vast configuration and policy combinations

Engineering Contradiction:
Improvereliability of provenance dataVSAvoidconfiguration and policy combinations
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the parameter of provenance entity priority from a static attribute to a dynamic parameter that can be adjusted based on system conditions. By introducing priority levels as a controllable parameter, the system simplifies configuration management while maintaining data reliability through selective retention policies.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamics by allowing the system to adaptively adjust provenance information retention based on storage capacity and priority levels. The system can dynamically move, compress, or delete provenance information according to current system state, reducing configuration complexity while preserving reliability.

Inventive Principle:
Principle #15Dynamics

3Productivity

If the system reduces storage for lower priority provenance information, then the storage efficiency is improved, but the loss of information increases

Engineering Contradiction:
Improvestorage efficiencyVSAvoidloss of provenance information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies discarding and recovering by moving low-priority provenance information to remote storage instead of permanently deleting it. This allows the system to improve local storage efficiency while preserving the ability to recover information when needed, balancing storage efficiency with information retention.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS12430068B2Managing provenance information for data processing pipelines
Publication Date: 2025.09.30 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US12430068B2 patent drawing
  • US12430068B2 patent drawing
  • US12430068B2 patent drawing

AI summary

A method for managing provenance information associated to one or more interconnected provenance entities in a provenance system for data processing pipelines in a distributed cloud environment over a network interface, wherein each of the data processing pipelines is configured to read in data, transform the data, and output transformed data is disclosed. The method comprises steps being performed by a configuration component of obtaining at least one declarative intent representing a configuration indicative of requirements and levels of priority for storage of provenance information for each of the data processing pipelines, deriving the requirements and levels of priority for storage of provenance information for each of the data processing pipelines based on the obtained at least one declarative intent, wherein one of the levels of priority—first level of priority—is higher than the other levels of priority—second levels of priority, estimating storage capacity for storage of provenance information in the provenance system based on the derived requirements and levels of priority, storing the provenance information according to the derived requirements and levels of priority for storage of provenance information and for each of the data processing pipelines, and when actual storage consumption for storage of provenance information in the provenance system meets a threshold of storage capacity set based on the estimated storage capacity: reducing a data amount for storage of provenance information of the second levels of priority in the provenance system. Corresponding computer program product, arrangement, configuration component, and system are also disclosed.