Resource Distribution via Data Lineage Attribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing resource distribution methods lack accuracy and efficiency in determining shares of resources to be allocated to data sources based on their knowledge contribution, often requiring data flow metering and manual reconciliation, which can slow processing and increase storage needs.
Innovation Solution
A method and system that receive values from multiple data sources, blend them into a dataset, assign data lineages, and determine knowledge contributions to allocate resources based on these contributions without measuring data flow, using attribute weights and priority scores to optimize resource distribution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data flow metering is used to determine resource shares, then resource distribution accuracy is improved, but processing complexity and storage requirements increase
Solution Approach 1:
The patent extracts the essential information needed for resource distribution (data lineages and knowledge contribution) from the complex data flow metering process. Instead of measuring actual data flow volumes, the system extracts metadata about data origins and relationships, eliminating the need for complex flow measurement infrastructure while maintaining distribution accuracy.
Solution Approach 2:
The patent introduces data lineages as an intermediary mechanism that mediates between data sources and resource distribution. Rather than directly measuring and tracking data flow between all components, the system uses data lineages to record provenance information, which then serves as the basis for determining knowledge contribution and resource allocation, simplifying the overall system architecture.
2Measurement precision
If manual reconciliation is used to determine knowledge contribution, then resource distribution accuracy is improved, but processing speed decreases
Solution Approach 1:
The patent performs preliminary action by automatically establishing and maintaining data lineages as data is ingested from various sources. This preliminary structuring of provenance information eliminates the need for subsequent manual reconciliation efforts, as the attribution framework is already in place when queries are executed, enabling both accuracy and speed.
Solution Approach 2:
The system implements self-service by automatically determining knowledge contribution through the pre-established data lineages. The attribution process serves itself by querying the existing lineage information rather than requiring external manual intervention for reconciliation, thereby maintaining accuracy while dramatically improving processing speed.
3Adaptability or versatility
If more data sources are handled, then data coverage is improved, but resource distribution accuracy decreases due to manual reconciliation limitations
Solution Approach 1:
The patent implements a universal data lineage framework that can handle multiple data sources with different formats, protocols, and characteristics through a single unified attribution mechanism. This multi-functional approach maintains resource distribution accuracy across diverse data sources by consistently applying the same provenance tracking and knowledge contribution calculation methods regardless of source type.
Data Source
AI summary
The present invention relates to a method for distributing resources to different data sources based on their knowledge contribution. In addition, the invention relates to a resource distribution data system for distributing resources, a computer program product for distributing resources and a computer readable medium. It may comprise the steps of receiving values from a plurality of different data sources, blending the received values for the attributes into a dataset, assigning data lineages to the values for the attributes, receiving a query for providing a data subset, providing the data subset based on the query, determining a knowledge contribution of each of the data sources to the data subset based on the data lineages of the values and instructing a distribution of shares of resources to the different data sources based on the knowledge contribution of each of the data sources to the data subset.


