Dataset Management System for Storage Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data warehouses face significant storage and computational resource challenges due to the large volumes of granular data transactions, which are not always necessary for analyses, leading to increased costs and resource consumption.

Innovation Solution

A dataset management system that determines a sampling rate based on the required level of accuracy, samples the data, and compares the sampled dataset with the full dataset to ensure sufficient accuracy, allowing for the deletion of the full dataset and storage of only the sampled dataset, thereby reducing data storage space and computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If all granular data transactions are stored in the data warehouse, then the completeness and accuracy of analysis results is improved, but the data storage space and costs increase significantly

Engineering Contradiction:
Improveaccuracy of analysis resultsVSAvoiddata storage space
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies partial action by storing only a sampled portion of data transactions rather than all granular data. The system determines an appropriate sampling rate that captures sufficient information for accurate analysis while discarding redundant data, thereby reducing storage space requirements while maintaining analysis accuracy within acceptable thresholds

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameter of data representation by transforming detailed granular transactions into aggregated or sampled data at different levels of granularity. By adjusting the sampling rate parameter, the system optimizes the balance between data completeness and storage efficiency, achieving accurate analysis results with reduced storage requirements

Inventive Principle:
Principle #35Parameter changes

2Reliability

If large volumes of granular data are stored for future analyses, then the reliability of analysis results is improved, but the computational resources and network resources consumption increase

Engineering Contradiction:
Improvereliability of analysis resultsVSAvoidcomputational resources consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system performs partial action by processing and storing only a representative sample of data transactions rather than the complete dataset. This sampling approach maintains the reliability of analysis results for most query types while significantly reducing the computational resources required for data processing, storage, and analysis operations

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If a high sampling rate is used to maintain data accuracy, then the accuracy of analysis results is improved, but the data storage space requirements increase

Engineering Contradiction:
Improveaccuracy of analysis resultsVSAvoiddata storage space
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies dynamics by making the sampling rate adjustable and adaptable rather than fixed. The system can dynamically determine and change the sampling rate based on specific analysis requirements, data characteristics, and storage constraints, allowing optimization of the balance between accuracy and storage space for different scenarios

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10545934B2Reducing data storage requirements
Publication Date: 2020.01.28 META PLATFORMS INC
  • US10545934B2 patent drawing
  • US10545934B2 patent drawing
  • US10545934B2 patent drawing

AI summary

A dataset management system (“system”) reduces the amount of data to be stored for future analyses. The system determines a sampling rate of the data based on a required level of accuracy, and samples the data at the determined sampling rate. Initially, all data transactions (“full dataset”) and the sampled data (“sampled dataset”) are logged and stored. Based upon a trigger condition, e.g., after a specified period, the full dataset and the sampled dataset are analyzed separately and the analysis results are compared. If the comparison is sufficiently similar (i.e., the sampling produces a sufficiently accurate set of data or a variance between the analysis results of the datasets is within a specified threshold), the system discontinues full data logging and stores only the sampled dataset. Further, the full dataset is deleted. The sampling thus reduces the required data volume significantly, thereby minimizing consumption of the storage space.