Synthesizing File Collections With Target Dedupability And Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Developers face challenges in simulating data streams generated by applications in developmental stages, as existing algorithms may not adequately meet the desired characteristics for various applications.
Innovation Solution
The system generates a collection of files based on data streams with desired parameters, using a filesystem synthesis module to create a simulated filesystem with specified folder/sub-folder structures, and populating it with data from the data stream to impart characteristics like dedupability, compression, clustering, and commonality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing algorithms are used to generate data streams, then data streams can be generated for testing, but they may not adequately meet the desired characteristics for various applications
Solution Approach 1:
The patent implements dynamic control of data stream characteristics through configurable parameters that allow real-time adjustment of dedupability, compression ratios, clustering patterns, and commonality. This enables the same generation system to adapt to different testing requirements while maintaining reliable characteristic delivery through structured parameter management.
Solution Approach 2:
The system employs multiple controllable parameters including dedupability parameters, compression parameters, clustering parameters, and commonality parameters to generate data streams with specific characteristics. By changing these parameters, the system can produce diverse data stream types (highly compressible, non-compressible, clustered, random) that reliably meet various application requirements.
2Reliability
If large data streams are processed directly for testing, then comprehensive testing can be performed, but it requires significant processing resources and time
Solution Approach 1:
The patent segments large data streams into smaller manageable units or blocks that can be processed independently and in parallel. This segmentation allows testing to be performed on representative samples while maintaining the statistical properties of the original data stream, significantly reducing processing time while preserving testing comprehensiveness.
Solution Approach 2:
The system generates synthetic data streams that copy the essential characteristics (dedupability, compression properties, clustering patterns) of real data streams without requiring actual large-scale data processing. These synthetic copies enable comprehensive characteristic testing with minimal processing time and resource consumption.
3Measurement precision
If data streams with specific characteristics are generated, then testing accuracy is improved, but the generation process becomes more complex
Solution Approach 1:
The patent implements a universal data stream generation platform that handles multiple characteristic types (dedupability, compression, clustering, commonality) through a single integrated system. This multi-functional approach reduces overall complexity by consolidating what would otherwise require separate generation tools for each characteristic type, while maintaining high testing accuracy through parameterized control.
Solution Approach 2:
The system achieves precise control of data stream characteristics through structured parameter management, where each characteristic (dedupability, compression ratio, clustering density, commonality level) is controlled by specific parameters. This parameterized approach simplifies the generation process compared to ad-hoc methods while enabling precise testing accuracy through systematic parameter adjustment.
Data Source
AI summary
One example method includes receiving a set of database parameters, creating one or more simulated databases based on the database parameters, receiving a set of target characteristics for the database, based on the target characteristics, slicing a datastream into a grouping of data slices, populating the simulated database(s) with the data slices to create the database collection and forward or reverse morphing the database from one generation to another without rewriting the entire database collection.


