Dynamic Peer Work Allocation for Storage Data Rebuild

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multi-device storage systems face bottlenecks in data recovery due to reliance on a storage control plane, leading to inefficiencies as the capacity and scalability of storage devices increase, particularly in peer-to-peer communication scenarios where rebuilding storage devices lack information about other devices' workloads.

Innovation Solution

Implementing a rebuild coordinator within each storage device to identify peer devices, determine work profiles, and order rebuild data units based on aggregate work factors, enabling dynamic work allocation and efficient peer-to-peer data rebuild operations without relying on a central control plane.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If storage devices rely on a storage control plane for data recovery operations, then data rebuild can be coordinated centrally, but bottlenecks occur and scalability deteriorates as the number and capacity of storage devices increase

Engineering Contradiction:
Improvedata recovery reliabilityVSAvoidrebuild operation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the centralized storage control plane functionality by distributing rebuild coordination capabilities to individual peer storage devices. Each storage device maintains local metadata and can independently coordinate rebuild operations with its peers, eliminating the single-point bottleneck while maintaining reliable data recovery through distributed consensus and coordination protocols.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If a centralized storage controller manages all storage devices in the array, then systematic control is achieved, but communication overhead and processing bottlenecks increase as system scale grows

Engineering Contradiction:
Improvestorage management controlVSAvoidcontrol plane complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts the rebuild coordination functionality from the centralized storage control plane and embeds it directly within peer storage devices. This allows storage devices to self-manage rebuild operations using local metadata and peer-to-peer communication, reducing control plane complexity and communication overhead while maintaining systematic control through distributed coordination.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If storage devices operate independently without knowledge of peer workloads, then device autonomy is maintained, but rebuild efficiency decreases due to inability to optimize data transfer scheduling

Engineering Contradiction:
Improvedevice autonomyVSAvoidrebuild operation speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements feedback mechanisms where storage devices exchange workload information and operational status with peers through metadata updates and status messages. This enables each device to adapt its rebuild operations based on real-time peer conditions, optimizing data transfer scheduling while maintaining device autonomy through decentralized decision-making based on shared feedback.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11182258B2Data rebuild using dynamic peer work allocation
Publication Date: 2021.11.23 WESTERN DIGITAL TECHNOLOGIES INC
  • US11182258B2 patent drawing
  • US11182258B2 patent drawing
  • US11182258B2 patent drawing

AI summary

Example peer storage systems, storage devices, and methods provide data rebuild across a peer communication channel using dynamic work allocation. A rebuild coordinator among the peer storage devices identifies peer storage devices including data units for the rebuild operation. The rebuild coordinator determines work profiles for the peer storage devices and uses the work profiles to determine the rebuild queue for the data units. The data is rebuilt according to the rebuild queue using the data units from the peer storage devices.