Centralized Data Library for Model Run Provenance Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data science teams face challenges in managing multiple versions of data and software components, leading to inefficiencies in data processing, collaboration, and storage, as they struggle to track the provenance and versions of input, intermediate, and output data across various stages of data science workflows.
Innovation Solution
A system and method that utilize a shared library to manage different versions of data science components, ensuring that the correct versions are used for processing, allowing for tracing the origin of resources, facilitating collaboration, and optimizing storage by keeping track of where copies are located and only creating them when necessary.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data science teams manage multiple versions of data and software components locally, then they can work with different data versions independently, but it leads to storage redundancy and difficulty in tracking provenance
Solution Approach 1:
The patent merges multiple local data storage locations into a centralized data library where all data versions are stored in a single location. This eliminates redundant storage across multiple workstations while maintaining version control through systematic organization and tracking mechanisms.
Solution Approach 2:
The centralized data library serves multiple functions: storing all data versions, tracking provenance information, enabling collaborative access, and providing a single source of truth for the entire data science team. This universal repository replaces multiple specialized local storage systems.
2Ease of operation
If data science teams create local copies of data for collaboration, then team members can work independently, but it increases storage costs and makes it difficult to ensure everyone uses the correct versions
Solution Approach 1:
The system performs preliminary actions by pre-organizing all data versions and provenance information in the centralized library before team members need them. This allows users to independently access any version they need without creating local copies, as everything is readily available in the central repository.
Solution Approach 2:
Team members can independently access and work with data versions through the centralized library without requiring others to share or update local copies. The system serves itself by maintaining an authoritative record of all versions and their relationships, eliminating the need for manual version synchronization.
3Quantity of substance
If data science teams store all data versions in a centralized library, then storage efficiency improves, but it becomes challenging to track the provenance and relationships between different data versions
Solution Approach 1:
The patent introduces provenance metadata as an intermediary layer that connects and describes relationships between different data versions in the centralized library. This metadata systematically records origin, transformations, and dependencies, enabling efficient storage while maintaining complete provenance information.
Data Source
AI summary
A system for tracking and representing data science model runs includes a hub including a first computing device communicatively coupled with a data store. A runner including a second computing device having a cache is communicatively coupled with the hub through a telecommunications network. An end user computing device includes a display and is communicatively coupled with the runner and the hub. User interfaces displayed on the display include: a unique identifier identifying a data science model run performed by the runner; a list of input files used by the runner to perform the run; a list of output files output by the runner as a result of the run; and a diagram diagramming a process flow including a visual representation of the input files, a visual representation of the run, and a visual representation of the output files.


