Configurable Volume Durability With SLA-Based Data Repair Timing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage systems face challenges in providing independently configurable durability for volumes with varying requirements, as they often allocate excess resources to higher durability needs while underprovisioning lower durability needs, and struggle to adapt to hardware anomalies and software bugs effectively.

Innovation Solution

A fault-tolerant data storage system with head nodes and data storage sleds that use translator components to determine target replacement times and allocate background bandwidth based on volume durability requirements, incorporating failure statistics and service level agreements to dynamically adjust resource allocation and ensure consistent performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data storage systems allocate resources uniformly to all volumes, then system simplicity is maintained, but resource efficiency deteriorates due to excess allocation to high durability volumes and underprovisioning of low durability volumes

Engineering Contradiction:
Improveresource efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements local quality by configuring durability parameters independently for each volume based on specific requirements. Each volume can have customized replication factors, erasure coding schemes, or durability thresholds, allowing resource allocation to match actual needs rather than applying a uniform policy across all storage volumes.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system applies dynamics by enabling runtime adjustment of durability configurations without requiring system reconfiguration or downtime. Durability parameters can be modified dynamically based on changing workload requirements, allowing the system to adapt resource allocation in real-time while maintaining operational continuity.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If data storage systems use fixed durability configurations, then system stability is maintained, but adaptability deteriorates when responding to hardware anomalies and software bugs

Engineering Contradiction:
Improveadaptability to hardware anomaliesVSAvoidsystem stability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system implements feedback mechanisms that continuously monitor storage system health, detecting hardware anomalies and software bugs in real-time. Based on this feedback, the system automatically adjusts durability configurations and triggers corrective actions such as increased replication or data recovery operations, maintaining stability while adapting to changing conditions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system applies preliminary action by pre-configuring multiple durability levels and recovery strategies that can be rapidly activated in response to detected anomalies. Before failures occur, the system establishes ready-to-execute recovery plans that can be immediately deployed when hardware or software issues are detected, reducing response time while maintaining system stability.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If data storage systems provide high durability guarantees, then data loss protection is improved, but resource consumption increases due to excessive replication and redundancy

Engineering Contradiction:
Improvedata durabilityVSAvoidstorage resource consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system implements parameter changes by offering a spectrum of durability configurations rather than a single fixed setting. Users can select from multiple replication factors, erasure coding ratios, or durability thresholds based on their specific requirements, allowing optimization between data protection levels and resource consumption without being constrained to maximum durability settings.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system applies partial action by implementing durability measures proportional to actual needs rather than applying maximum protection uniformly. Critical data receives higher durability guarantees with full replication, while less critical data uses minimal redundancy or erasure coding, optimizing resource usage by applying only the necessary level of protection for each data type.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11853587B2Data storage system with configurable durability
Publication Date: 2023.12.26 AMAZON TECH INC
  • US11853587B2 patent drawing
  • US11853587B2 patent drawing
  • US11853587B2 patent drawing

AI summary

A fault-tolerant data storage system associates durability requirements of service level agreements (SLAs) for volumes stored in the fault-tolerant data storage system with volume partitions stored in the fault-tolerant data storage system. For a given volume partition, volume data is stored in two or more replicas on two or more different system components and/or erasure encoded across multiple other system components. The fault-tolerant data storage system uses the respective durability requirements of the SLAs and failure statistics of the system components to allocate bandwidth for replacing lost instances of redundantly stored volume data such that the lost data is replaced within a target time calculated to guarantee the durability requirements of the SLAs are satisfied.