De-clustered Disk Array I/O Scheduling for RAID Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Long data reconstruction times in disk arrays increase the risk of data loss due to higher probabilities of multiple physical disk failures during reconstruction, especially with RAID-6 data protection algorithms, which can lead to inefficient distribution of I/O requests and reduced overall performance.

Innovation Solution

A de-clustered disk array architecture is implemented, where physical disks are divided into schedule groups based on active I/O requests, with a priority-based scheduling system to distribute I/O requests evenly, and a virtual storage solution that allows all working disks to participate in data reconstruction, reducing data reconstruction time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Volume of stationary object

If data is stored in a traditional disk array with growing disk capacity, then storage capacity increases, but data reconstruction time becomes longer

Engineering Contradiction:
Improvestorage capacityVSAvoiddata reconstruction time
Core Design Contradiction:
Volume of stationary objectVSLoss of time

Solution Approach 1:

The patent segments the disk array into multiple groups (first group and second group) with different scheduling priorities. This segmentation allows parallel reconstruction operations across multiple disks simultaneously, reducing overall reconstruction time while maintaining large storage capacity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic scheduling that adjusts disk selection based on real-time I/O statistics and workload conditions. The scheduler dynamically changes which disks are used for reconstruction versus normal I/O operations, optimizing reconstruction speed without compromising storage capacity.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If traditional scheduling is used in de-clustered disk arrays, then system simplicity is maintained, but I/O requests concentrate on specific disks reducing performance

Engineering Contradiction:
Improvescheduling system complexityVSAvoidI/O performance
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent implements a feedback mechanism where the scheduler continuously monitors I/O statistics from each disk and adjusts scheduling decisions accordingly. Disks with lower current workload receive higher priority for new I/O requests, automatically balancing load across the array and preventing concentration on specific disks.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the scheduling parameter from static to dynamic by incorporating real-time I/O statistics. The scheduler modifies disk selection based on current workload parameters, achieving balanced I/O distribution and improved performance without significantly increasing system complexity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If RAID-6 data protection algorithm is used, then data protection capability is improved, but data reconstruction time increases leading to higher data loss risk

Engineering Contradiction:
Improvedata protection capabilityVSAvoiddata reconstruction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments reconstruction operations across multiple disks in different groups, allowing parallel processing of reconstruction tasks. This maintains RAID-6's strong data protection capability while reducing overall reconstruction time by distributing the computational load across multiple disks simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-selecting disks for reconstruction based on current I/O statistics before reconstruction actually begins. This preliminary scheduling ensures that disks with lower current workload are chosen, minimizing the impact on ongoing operations and reducing effective reconstruction time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9798471B2Performance of de-clustered disk array by disk grouping based on I/O statistics
Publication Date: 2017.10.24 EMC IP HLDG CO LLC
  • US9798471B2 patent drawing
  • US9798471B2 patent drawing
  • US9798471B2 patent drawing

AI summary

Embodiments of the present invention relate to a method and apparatus for improving performance of a de-clustered disk array by making statistics on a number and types of active input/output (I/O) requests of each of the plurality of physical disks; dividing the plurality of physical disks at least into a first schedule group and a second schedule group based on the statistic number and types of the active I/O requests of the each physical disk for a predetermined time period, the first schedule group having a first schedule priority, the second schedule group having a second schedule priority higher than the first schedule priority; and selecting, in a decreasing order of the schedule priority, a physical disk for schedule from one of the resulting schedule groups thereby preventing too many I/O requests from concentrating on some physical disks and thereby improve overall performance of a de-clustered RAID.