Provenance Storage Prioritization for Distributed Data Pipelines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Provenance systems in distributed cloud environments face challenges in managing vast arrays of configuration and policy combinations, leading to cumbersome handling and inefficient storage of provenance information, which is crucial for data processing pipelines.
Innovation Solution
A method and system that derive declarative intents representing configuration requirements and priority levels for storing provenance information, estimate storage capacity, and reduce data amount by compressing, removing indexing, or moving lower priority data to remote storage when thresholds are met, ensuring efficient resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If provenance systems capture comprehensive provenance information for data processing pipelines, then the completeness and usefulness of provenance data is improved, but the storage capacity consumption increases significantly
Solution Approach 1:
The patent applies local quality by assigning different retention policies to different provenance entities based on their priority levels. High-priority entities maintain complete provenance information locally, while low-priority entities have their information moved to remote storage or compressed, thereby optimizing storage usage while preserving information quality where needed.
Solution Approach 2:
The patent segments provenance information management into priority-based categories. By dividing provenance entities into high-priority and low-priority groups, the system can apply different storage strategies to each segment, reducing overall storage consumption while maintaining essential information.
2Reliability
If the system stores detailed provenance information for all data processing pipelines, then the reliability of provenance data is improved, but the device complexity increases due to vast configuration and policy combinations
Solution Approach 1:
The patent changes the parameter of provenance entity priority from a static attribute to a dynamic parameter that can be adjusted based on system conditions. By introducing priority levels as a controllable parameter, the system simplifies configuration management while maintaining data reliability through selective retention policies.
Solution Approach 2:
The patent introduces dynamics by allowing the system to adaptively adjust provenance information retention based on storage capacity and priority levels. The system can dynamically move, compress, or delete provenance information according to current system state, reducing configuration complexity while preserving reliability.
3Productivity
If the system reduces storage for lower priority provenance information, then the storage efficiency is improved, but the loss of information increases
Solution Approach 1:
The patent applies discarding and recovering by moving low-priority provenance information to remote storage instead of permanently deleting it. This allows the system to improve local storage efficiency while preserving the ability to recover information when needed, balancing storage efficiency with information retention.
Data Source
AI summary
A method for managing provenance information associated to one or more interconnected provenance entities in a provenance system for data processing pipelines in a distributed cloud environment over a network interface, wherein each of the data processing pipelines is configured to read in data, transform the data, and output transformed data is disclosed. The method comprises steps being performed by a configuration component of obtaining at least one declarative intent representing a configuration indicative of requirements and levels of priority for storage of provenance information for each of the data processing pipelines, deriving the requirements and levels of priority for storage of provenance information for each of the data processing pipelines based on the obtained at least one declarative intent, wherein one of the levels of priority—first level of priority—is higher than the other levels of priority—second levels of priority, estimating storage capacity for storage of provenance information in the provenance system based on the derived requirements and levels of priority, storing the provenance information according to the derived requirements and levels of priority for storage of provenance information and for each of the data processing pipelines, and when actual storage consumption for storage of provenance information in the provenance system meets a threshold of storage capacity set based on the estimated storage capacity: reducing a data amount for storage of provenance information of the second levels of priority in the provenance system. Corresponding computer program product, arrangement, configuration component, and system are also disclosed.


