Spare Disk Drives Overprovisioning RAID Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data storage systems inefficiently utilize spare disk drives, which consume resources without actively contributing to performance or redundancy until a disk drive failure occurs.

Innovation Solution

Transferring data segments from operating RAID group disk drives to spare regions on spare disk drives to create unused space, thereby overprovisioning storage and distributing workload, and rebuilding data upon a disk drive failure using both segment data and data from functioning drives.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If spare disk drives are kept in a powered state ready to replace failed drives, then reliability is improved by enabling quick replacement, but device complexity and resource consumption increase as drives occupy space and consume power without actively contributing to performance

Engineering Contradiction:
ImprovereliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The spare disk drives are configured to perform dual functions: serving as hot spares for immediate failure replacement and providing overprovisioning for workload distribution. The system dynamically assigns drives between these roles based on operational needs, allowing the same physical drives to contribute to both reliability and performance optimization without requiring separate dedicated spares

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements dynamic role assignment where spare drives can be flexibly allocated as overprovisioned storage when no failures occur, and automatically switched to failure replacement mode when needed. This dynamic reconfiguration optimizes resource utilization by adapting the function of spare drives to current system conditions

Inventive Principle:
Principle #15Dynamics

2Reliability

If spare disk drives are used solely for failure replacement, then reliability is maintained through hot spares, but productivity decreases as drives consume power and occupy space without actively contributing to storage performance

Engineering Contradiction:
ImprovereliabilityVSAvoidproductivity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Spare drives perform multiple functions by serving as both hot spares for failure replacement and as overprovisioned storage for active workload distribution, eliminating the idle resource problem while maintaining reliability capabilities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system ensures continuous useful action from spare drives by having them actively participate in workload distribution during normal operation, rather than remaining idle. This continuous utilization improves overall system productivity while reliability is maintained through the ability to quickly swap drives when failures occur

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If data is transferred to spare regions on spare disk drives to overprovision storage, then productivity is improved by distributing workload, but device complexity increases due to the need to manage and track spare regions

Engineering Contradiction:
ImproveproductivityVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system introduces a mapping layer that intermediates between logical storage addresses and physical locations on drives including spare regions. This mapping mechanism simplifies the management of complex spare region configurations by abstracting the details of data distribution across traditional and spare drives

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The storage system is segmented into traditional RAID groups and overprovisioning groups utilizing spare drives. This segmentation allows independent management of each group type while maintaining unified system control, reducing the complexity of managing mixed-purpose drives

Inventive Principle:
Principle #1Segmentation

4Duration of action of stationary object

If SSDs are overprovisioned by transferring data segments to spare regions, then the life expectancy of SSDs is extended by reducing write amplification, but loss of time occurs during the data transfer process

Engineering Contradiction:
Improvelife expectancyVSAvoidloss of time
Core Design Contradiction:
Duration of action of stationary objectVSLoss of time

Solution Approach 1:

The system performs preliminary data transfer to spare regions during low-utilization periods or maintenance windows, preparing the overprovisioning structure in advance. This preliminary action allows the overprovisioning benefits to be realized without causing time loss during peak operational periods when the SSDs are actively serving workloads

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9921912B1Using spare disk drives to overprovision raid groups
Publication Date: 2018.03.20 EMC IP HLDG CO LLC
  • US9921912B1 patent drawing
  • US9921912B1 patent drawing
  • US9921912B1 patent drawing

AI summary

A technique for managing spare disk drives in a data storage system includes transferring segments of data from disk drives of an operating RAID group to spare regions on a set of spare disk drives to create unused space in the disk drives of the RAID group, thus using the spare regions to overprovision storage in the RAID group. Upon a failure of one of the disk drives in the RAID group, data of the failing disk drive are rebuilt based on the segments of data as well as on data from still-functioning disk drives in the RAID group. Thus, the spare disk drives act not only to overprovision storage for the RAID group prior to disk drive failure, but also to fulfill their role as spares in the event of a disk drive failure.