Data Value Tree for Analytics Provenance Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analytics systems lack mechanisms to measure the value of data science efforts, associate actual and predicted values with data science experiments, and distribute value among contributing sources, tools, and personnel, leading to lost provenance and inability to evaluate tool and personnel contributions.
Innovation Solution
A data value tree is created and stored in a catalog during data science experiments, allowing for the assignment and propagation of value across data structure elements, enabling the tracking and querying of value contributions from data, tools, and personnel, and facilitating the correlation of value across projects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data analytics systems execute analytic processes on large data sets, then valuable insights and recommendations are generated, but the value and provenance of contributions from data, tools, and personnel are lost
Solution Approach 1:
The patent segments the analytic process into distinct components (data sources, tools, personnel, intermediate results, final recommendations) and creates separate data structure elements for each. This segmentation allows tracking of individual contributions while maintaining the overall analytic workflow, resolving the contradiction by preserving provenance information without hindering productivity.
Solution Approach 2:
The patent introduces an intermediary data structure (the catalog with data structure elements) that mediates between the analytic process execution and the value attribution. This intermediary captures and stores provenance information about contributions from data, tools, and personnel, enabling value tracking without interfering with the analytic productivity.
2Adaptability or versatility
If companies make large investments in data, data scientists, and analytics tools, then analytic capabilities are enhanced, but the ability to measure and evaluate the value of these investments is lacking
Solution Approach 1:
The patent changes the parameter space by introducing value attribution parameters to existing data structure elements. Each element (data source, tool, personnel contribution) is assigned value parameters that quantify its contribution to the analytic outcomes. This enables precise measurement of investment value while maintaining enhanced analytic capabilities.
3Productivity
If analytic processes are executed within an analytic computing environment, then data science experiments are performed, but there is no mechanism to associate actual and predicted values with the experiments
Solution Approach 1:
The patent applies preliminary action by creating data structure elements and assigning value parameters before the analytic process completes. The system prepares the infrastructure for value tracking in advance, associating predicted values with data structure elements during execution, and then updating with actual values afterward. This preliminary preparation enables seamless value association without hindering experiment execution productivity.
Data Source
AI summary
At least part of an analytic process is executed on one or more data sets. Execution of the analytic process is performed within an analytic computing environment. During the course of execution of the analytic process, a data structure is generated comprising data structure elements. The data structure elements represent attributes associated with execution of the analytic process. Value is assigned to at least a portion of the data structure elements. The data structure generated during execution of the analytic process may be stored in an accessible catalog of other data structures generated during execution of other analytic processes.


