Data Set Distribution via Segmentation and Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing volume of big data exceeds the processing capabilities of conventional software tools, making it difficult for entities to manage and analyze large data sets effectively, particularly in industries like finance where fine-grained data is collected.

Innovation Solution

A method is described to divide large data sets into smaller data files, which are then stored in distributed data stores and made available through a data marketplace, allowing customers to acquire and process specific subsets using metadata for identification and electronic signatures for access control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If large data sets are stored and managed using conventional software tools, then data storage capacity is sufficient, but processing time and computational resources become excessive

Engineering Contradiction:
Improvedata storage capacityVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent divides large data sets into smaller, manageable data files that can be processed independently. This segmentation allows conventional software tools to handle each file efficiently without being overwhelmed by the entire large data set, thereby reducing processing time while maintaining the ability to store substantial quantities of data.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If fine-grained data is captured at high granularity levels, then data detail and analysis precision improve, but data volume and management complexity increase

Engineering Contradiction:
Improvedata granularityVSAvoiddata management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments high-granularity data into individual data files that can be managed separately. This allows the system to maintain fine-grained data detail for analysis precision while reducing management complexity by breaking down the large data set into smaller, more manageable units that can be processed and organized more efficiently.

Inventive Principle:
Principle #1Segmentation

3Productivity

If data is made accessible to multiple customers through a marketplace, then data utilization and productivity improve, but access control and security requirements increase

Engineering Contradiction:
Improvedata utilizationVSAvoidaccess control security
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent creates copies of data files that can be distributed to multiple customers through the marketplace. Each customer receives a copy that can be processed independently, enabling simultaneous data utilization by multiple parties while maintaining security through controlled distribution mechanisms.

Inventive Principle:
Principle #26Copying

4Productivity

If data sets are divided into multiple data files, then data processing efficiency and storage needs are reduced, but data organization and retrieval complexity increase

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddata organization complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a data marketplace as an intermediary system that manages the organization and retrieval of divided data files. This marketplace provides standardized interfaces and mechanisms for customers to search, select, and access data files, thereby reducing the complexity that would otherwise result from managing fragmented data structures.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10108692B1Data set distribution
Publication Date: 2018.10.23 AMAZON TECH INC
  • US10108692B1 patent drawing
  • US10108692B1 patent drawing
  • US10108692B1 patent drawing

AI summary

A method is described for distributing a data set. The method may include dividing a data set into a number of data groupings based on a data set attribute value. The groupings of data may be stored in a data store and may be associated with metadata that describes a grouping of data. A grouping of data may be distributed by generating a reference that may be used to access the grouping of data in the data store. The reference may include information that enables access to the grouping of data. When presented, the information included in the reference may be authenticated whereupon the grouping of data may be provided.