Distributed Data Store for Scalable Data Center Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data center testing systems are inadequate in simulating complex workflows, scaling to millions of files, supporting integrity checking, and managing resources efficiently, often degrading performance and risking customer data due to storage constraints and resource limitations.

Innovation Solution

A distributed data store system using fault-tolerant, in-memory databases across test clients, with an orchestrator component for managing IO operations, snapshot data, and integrity verification, allowing for scalable and realistic testing without impacting actual data center performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If testing data is stored on the primary data center storage system, then testing can be performed, but data center performance degrades and customer data security is compromised

Engineering Contradiction:
Improvetesting capabilityVSAvoiddata center performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system segments the storage infrastructure into two independent parts: primary storage for customer data and secondary storage for testing data. This segmentation allows testing operations to occur on isolated storage resources, preventing performance degradation of the primary data center system while maintaining full testing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A secondary storage system acts as an intermediary between the testing workflow and the primary storage system. This intermediary provides the necessary storage capacity for testing without directly impacting the primary system, enabling comprehensive testing while protecting customer data and maintaining performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If storage capacity is increased to support large-scale testing, then more files can be tested, but resource consumption increases and costs rise

Engineering Contradiction:
Improvenumber of test filesVSAvoidresource consumption
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

Instead of creating entirely new storage infrastructure, the system creates a secondary storage environment that copies or mirrors the essential characteristics of the primary storage system. This allows testing of millions of files with realistic data volumes while consuming resources only at the secondary level, avoiding excessive resource consumption.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the storage parameter from shared primary storage to dedicated secondary storage, fundamentally altering how storage resources are allocated and consumed. This parameter change enables large-scale testing with millions of files while keeping resource consumption isolated to non-critical systems.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If testing workflows are simplified, then testing can be performed quickly, but complex real-world scenarios cannot be simulated

Engineering Contradiction:
Improvetesting speedVSAvoidworkflow complexity
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The testing system implements dynamic workflow capabilities that can adapt to varying levels of complexity based on testing needs. The orchestrator component dynamically manages and coordinates complex multi-step workflows involving data retrieval, processing, analysis, and reporting, allowing the system to handle anything from simple to highly complex testing scenarios without sacrificing speed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where the orchestrator monitors workflow execution and adjusts resource allocation, data retrieval strategies, and processing priorities in real-time. This feedback loop enables the system to maintain high productivity while handling complex workflows by optimizing resource usage based on actual testing requirements.

Inventive Principle:
Principle #23Feedback

4Duration of action of moving object

If testing duration is extended to cover prolonged scenarios, then more comprehensive testing is achieved, but system resources are depleted

Engineering Contradiction:
Improvetesting durationVSAvoidresource availability
Core Design Contradiction:
Duration of action of moving objectVSLoss of energy

Solution Approach 1:

The system segments testing operations into independent workflows that can run in parallel on the secondary storage system. This segmentation allows prolonged and comprehensive testing to occur without depleting primary system resources, as each workflow operates independently with its own resource allocation on the isolated secondary infrastructure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The secondary storage system and orchestrator are designed as universal testing infrastructure that can handle multiple different testing scenarios simultaneously. This multi-functionality allows the system to run diverse, prolonged testing workflows without resource depletion, as the universal platform efficiently manages and shares resources across all testing activities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11500749B2Distributed data store for testing data center services
Publication Date: 2022.11.15 EMC IP HLDG CO LLC
  • US11500749B2 patent drawing
  • US11500749B2 patent drawing
  • US11500749B2 patent drawing

AI summary

Architectures and techniques are described that can enhance or improve testing procedure that tests operation of services provided by a data center. Advantageously, the testing dataset can be distributed on test clients allowing the testing procedure to scale to any suitable size, while providing integrity checking for a dataset that includes snapshots and scalability to millions of files, while supporting multiple readers/writers.