Centralized Data Library for Model Run Provenance Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data science teams face challenges in managing multiple versions of data and software components, leading to inefficiencies in data processing, collaboration, and storage, as they struggle to track the provenance and versions of input, intermediate, and output data across various stages of data science workflows.

Innovation Solution

A system and method that utilize a shared library to manage different versions of data science components, ensuring that the correct versions are used for processing, allowing for tracing the origin of resources, facilitating collaboration, and optimizing storage by keeping track of where copies are located and only creating them when necessary.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data science teams manage multiple versions of data and software components locally, then they can work with different data versions independently, but it leads to storage redundancy and difficulty in tracking provenance

Engineering Contradiction:
Improveability to work with different data versionsVSAvoidstorage redundancy
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges multiple local data storage locations into a centralized data library where all data versions are stored in a single location. This eliminates redundant storage across multiple workstations while maintaining version control through systematic organization and tracking mechanisms.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The centralized data library serves multiple functions: storing all data versions, tracking provenance information, enabling collaborative access, and providing a single source of truth for the entire data science team. This universal repository replaces multiple specialized local storage systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If data science teams create local copies of data for collaboration, then team members can work independently, but it increases storage costs and makes it difficult to ensure everyone uses the correct versions

Engineering Contradiction:
Improveindependent work capabilityVSAvoidstorage costs
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The system performs preliminary actions by pre-organizing all data versions and provenance information in the centralized library before team members need them. This allows users to independently access any version they need without creating local copies, as everything is readily available in the central repository.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Team members can independently access and work with data versions through the centralized library without requiring others to share or update local copies. The system serves itself by maintaining an authoritative record of all versions and their relationships, eliminating the need for manual version synchronization.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If data science teams store all data versions in a centralized library, then storage efficiency improves, but it becomes challenging to track the provenance and relationships between different data versions

Engineering Contradiction:
Improvestorage efficiencyVSAvoidprovenance tracking
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent introduces provenance metadata as an intermediary layer that connects and describes relationships between different data versions in the centralized library. This metadata systematically records origin, transformations, and dependencies, enabling efficient storage while maintaining complete provenance information.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11829853B2Systems and methods for tracking and representing data science model runs
Publication Date: 2023.11.28 DATADIRECT NETWORKS INC
  • US11829853B2 patent drawing
  • US11829853B2 patent drawing
  • US11829853B2 patent drawing

AI summary

A system for tracking and representing data science model runs includes a hub including a first computing device communicatively coupled with a data store. A runner including a second computing device having a cache is communicatively coupled with the hub through a telecommunications network. An end user computing device includes a display and is communicatively coupled with the runner and the hub. User interfaces displayed on the display include: a unique identifier identifying a data science model run performed by the runner; a list of input files used by the runner to perform the run; a list of output files output by the runner as a result of the run; and a diagram diagramming a process flow including a visual representation of the input files, a visual representation of the run, and a visual representation of the output files.