Synthetic Data Generation for Secure Storage Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating deduplication and compression systems in data storage are time-consuming and expose confidential information, making it difficult for entities to compare vendor designs efficiently without compromising data security.
Innovation Solution
Generating synthetic data sets with the same deduplication and compression characteristics as existing data, using hash values, random substitution ciphers, and shuffled run lengths to create anonymized data that can be evaluated by vendors without exposing PII, ensuring accurate and valid comparisons while protecting sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real data is used for evaluation, then evaluation accuracy is improved, but data security is compromised
Solution Approach 1:
The patent creates synthetic copies of real data that preserve the statistical properties, patterns, and characteristics needed for accurate evaluation, while containing no actual confidential information. These synthetic datasets replicate the structure and behavior of real data without exposing sensitive content.
Solution Approach 2:
The patent introduces synthetic data as an intermediary between the need for accurate evaluation and the requirement for data security. This intermediary preserves evaluation validity while eliminating security risks associated with using real confidential data.
2Object-affected harmful factors
If traditional evaluation process is used, then data security is maintained, but evaluation time increases
Solution Approach 1:
The patent performs preliminary actions by generating synthetic datasets in advance that can be freely distributed to vendors for evaluation. This eliminates the need for lengthy coordination and data transfer processes while maintaining security, as synthetic data can be shared without confidentiality concerns.
3Reliability
If real data is shared with vendors, then evaluation validity is improved, but information exposure risk increases
Solution Approach 1:
The patent creates faithful replicas of real data structures and patterns that enable valid evaluation of deduplication and compression systems, while ensuring no actual confidential information is exposed. The synthetic copies maintain all necessary statistical properties for accurate testing.
Data Source
AI summary
Deduplication and compression evaluation methods and systems involve one or more processors obfuscating plain text file data in each file of a computer file system using a first cipher encryption scheme, obfuscating each plain text file name representing the plain text file data in each file of the computer file system using a second cipher encryption scheme, and associating each obfuscated file name representing the plain text file data of each of the plurality of files of the computer file system with the obfuscated file data of each of the plurality of files of the computer system. In addition, each plain text directory name for each of the obfuscated file names associated with the obfuscated file data in each of the plurality of files of the computer file system is obfuscated using a third cipher encryption scheme.


