Provenance Dependency Deriver for Data Stream Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In data stream processing systems, manually obtaining provenance reports is inefficient, as it lacks automation and accuracy in tracking the origins and justification of events, especially in high-volume data applications like healthcare, where timely and reliable information is crucial for medical decisions.

Innovation Solution

The implementation of a Provenance Dependency Deriver (PDD) component that automatically generates and validates dependency equations by observing input and output data streams, using techniques like principal component analysis, information-theoretic constructs, and empirical probability distributions to establish time-invariant temporal dependency functions, thereby identifying which input elements contribute to specific output elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If manual provenance reports are obtained by leveraging developer specifications, then the system can track event origins, but the process is inefficient and lacks automation

Engineering Contradiction:
Improveautomation of provenance determinationVSAvoidefficiency of provenance report generation
Core Design Contradiction:
Extent of automationVSProductivity

Solution Approach 1:

The system automatically determines provenance information by observing input-output data associations without requiring manual intervention or developer specifications. The Provenance Dependency Deriver component self-determines dependency equations by analyzing actual data flow patterns, eliminating the need for manual provenance report generation while improving both automation extent and productivity

Inventive Principle:
Principle #25Self-service

2Measurement precision

If dependency equations are automatically determined by observing data associations, then automation and accuracy improve, but system complexity increases

Engineering Contradiction:
Improveaccuracy of provenance trackingVSAvoidcomplexity of dependency determination mechanism
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The Provenance Dependency Deriver component acts as an intermediary between the data stream processing elements and the provenance tracking system. It observes input-output associations and computes dependency equations as intermediate representations, providing accurate provenance tracking while encapsulating the complexity within a dedicated component that can be independently managed and validated

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If provenance data is validated by computing and comparing intervals, then reliability improves, but processing time increases

Engineering Contradiction:
Improvereliability of provenance validationVSAvoidtime required for interval validation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary computation of dependency equations and interval specifications during normal data stream processing, rather than performing validation separately afterward. By continuously observing input-output associations and maintaining updated dependency models, the system prepares validation information in advance, reducing the time required for explicit validation while maintaining high reliability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8775344B2Determining and validating provenance data in data stream processing system
Publication Date: 2014.07.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8775344B2 patent drawing
  • US8775344B2 patent drawing
  • US8775344B2 patent drawing

AI summary

Techniques are disclosed for determining and validating provenance data in such data stream processing systems. For example, a method for processing data associated with a data stream received by a data stream processing system, wherein the system comprises a plurality of processing elements, comprises the following steps. Input data elements and output data elements associated with at least one processing element of the plurality of processing elements are obtained. One or more intervals are computed for the processing element using data representing observations of associations between inputs elements and output elements of the processing element, wherein, for a given one of the intervals, one or more particular input elements contained within the given interval are determined to have contributed to a particular output element. In another method, intervals are specified, and then validated by comparing the specified intervals against intervals computed based on observations.