Data Set Distribution via Segmentation and Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing volume of big data exceeds the processing capabilities of conventional software tools, making it difficult for entities to manage and analyze large data sets effectively, particularly in industries like finance where fine-grained data is collected.
Innovation Solution
A method is described to divide large data sets into smaller data files, which are then stored in distributed data stores and made available through a data marketplace, allowing customers to acquire and process specific subsets using metadata for identification and electronic signatures for access control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If large data sets are stored and managed using conventional software tools, then data storage capacity is sufficient, but processing time and computational resources become excessive
Solution Approach 1:
The patent divides large data sets into smaller, manageable data files that can be processed independently. This segmentation allows conventional software tools to handle each file efficiently without being overwhelmed by the entire large data set, thereby reducing processing time while maintaining the ability to store substantial quantities of data.
2Measurement precision
If fine-grained data is captured at high granularity levels, then data detail and analysis precision improve, but data volume and management complexity increase
Solution Approach 1:
The patent segments high-granularity data into individual data files that can be managed separately. This allows the system to maintain fine-grained data detail for analysis precision while reducing management complexity by breaking down the large data set into smaller, more manageable units that can be processed and organized more efficiently.
3Productivity
If data is made accessible to multiple customers through a marketplace, then data utilization and productivity improve, but access control and security requirements increase
Solution Approach 1:
The patent creates copies of data files that can be distributed to multiple customers through the marketplace. Each customer receives a copy that can be processed independently, enabling simultaneous data utilization by multiple parties while maintaining security through controlled distribution mechanisms.
4Productivity
If data sets are divided into multiple data files, then data processing efficiency and storage needs are reduced, but data organization and retrieval complexity increase
Solution Approach 1:
The patent introduces a data marketplace as an intermediary system that manages the organization and retrieval of divided data files. This marketplace provides standardized interfaces and mechanisms for customers to search, select, and access data files, thereby reducing the complexity that would otherwise result from managing fragmented data structures.
Data Source
AI summary
A method is described for distributing a data set. The method may include dividing a data set into a number of data groupings based on a data set attribute value. The groupings of data may be stored in a data store and may be associated with metadata that describes a grouping of data. A grouping of data may be distributed by generating a reference that may be used to access the grouping of data in the data store. The reference may include information that enables access to the grouping of data. When presented, the information included in the reference may be authenticated whereupon the grouping of data may be provided.


