Synthetic Data Generation for Block Storage Backup Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage systems face challenges in efficiently simulating real-world data manipulation patterns for block-based storage systems, which is essential for effective backup and testing, as they lack synthetic data that accurately parallels real-world data changes.

Innovation Solution

A synthetic data generation system that mimics real-world data modification patterns by selecting and modifying tracks and blocks in a manner analogous to real-world client interactions, using a synthetic data generation client that communicates with a protection storage system to create a dataset reflecting up-to-date changes, thereby simulating the data protection process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real-world data is used for backup testing, then testing accuracy is improved, but data transfer time and storage resources are increased

Engineering Contradiction:
Improvetesting accuracyVSAvoiddata transfer time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic data that copies the essential characteristics and patterns of real-world data without transferring actual data. The synthetic data generation system replicates data modification patterns, access patterns, and structural properties, enabling accurate backup system testing while avoiding the time and resource costs of copying real data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary analysis of real-world data patterns to establish synthetic data generation rules before actual testing begins. By pre-characterizing data modification patterns, access sequences, and structural properties, the system prepares synthetic data templates that can quickly generate realistic test scenarios without requiring real data transfer during testing.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If real-world data is used for backup testing, then testing accuracy is improved, but storage resources are increased

Engineering Contradiction:
Improvetesting accuracyVSAvoidstorage resources
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates synthetic data that copies the essential characteristics and patterns of real-world data without transferring actual data. The synthetic data generation system replicates data modification patterns, access patterns, and structural properties, enabling accurate backup system testing while avoiding the time and resource costs of copying real data.

Inventive Principle:
Principle #26Copying

3Productivity

If synthetic data is used to parallel real-world data changes, then backup efficiency is improved, but data generation complexity is increased

Engineering Contradiction:
Improvebackup efficiencyVSAvoiddata generation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transforms complex data pattern recognition into simplified parameter representations. By extracting key characteristics such as modification frequency, access patterns, and structural properties as quantifiable parameters, the system can generate synthetic data using controlled parameter adjustments rather than complex rule-based systems, reducing generation complexity while maintaining realism.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9128823B1Synthetic data generation for backups of block-based storage
Publication Date: 2015.09.08 EMC IP HLDG CO LLC
  • US9128823B1 patent drawing
  • US9128823B1 patent drawing
  • US9128823B1 patent drawing

AI summary

A system and method for generating synthetic data to simulate backing up data between a primary storage system and a protection storage system is presented. In one embodiment, a first track in a set of tracks is selected at random. Having selected a first track, at least a first block in the first track is modified. Subsequently, it is determined, based on a track run probability, whether to modify a second track that is consecutive to the first track or a third track that is selected randomly. Depending on the determination, at least one block is modified at either the second or third track. Other embodiments are also described herein.