Multi-drive sled servicing via lock-step isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multi-carrier drive facilities face challenges in achieving high drive density while minimizing the impact of service operations, as retracting a sled to access a failed drive decouples all storage devices on that sled, including non-faulty ones, often requiring the entire storage array to be brought offline.

Innovation Solution

Implementing a lock-step process between the operating system and hardware elements, with a user interface that indicates when it is safe to begin service, allowing for selective decoupling of faulty drives and enabling preparations for service without affecting operational drives, ensuring minimal disruption to the storage facility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of repair

If a sled is retracted to access a failed drive in a multi-carrier drive facility, then service accessibility is improved, but all storage devices on that sled are decoupled from the storage array including non-faulty ones, reducing productivity

Engineering Contradiction:
Improveservice accessibilityVSAvoidstorage array availability
Core Design Contradiction:
Ease of repairVSProductivity

Solution Approach 1:

The system segments the decoupling operation to affect only the specific faulty drive rather than the entire sled. The lock-step process enables individual drive isolation while keeping other drives on the same sled connected to the storage array, thus maintaining productivity while enabling service access.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by establishing a lock-step relationship between the operating system and hardware elements before service begins. This preliminary coordination allows the system to prepare for selective decoupling, ensuring that only necessary components are disconnected while others remain operational.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the entire storage array is brought offline for service operations, then data integrity is protected, but service time and productivity are reduced

Engineering Contradiction:
Improvedata integrityVSAvoidservice time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system applies local quality by implementing service isolation at the drive level rather than array level. The lock-step mechanism enables different states for different drives - the faulty drive is decoupled for service while other drives remain coupled and operational, thus protecting data integrity locally without requiring global offline status.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of bringing down the entire storage array (excessive action), the system performs partial action by isolating only the specific faulty drive. This partial decoupling approach maintains sufficient data integrity protection while minimizing service time and productivity loss.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If multiple storage devices are housed on a single sled to increase drive density, then space efficiency is improved, but service complexity increases as all devices must be accessed together

Engineering Contradiction:
Improvedrive densityVSAvoidservice complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the service control mechanism to operate at the individual drive level within the sled architecture. The lock-step process creates independent control paths that allow service personnel to address specific drives without affecting others on the same sled, thus maintaining high drive density while reducing service complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10929019B1Enhanced servicing for multi-drive sleds
Publication Date: 2021.02.23 EMC IP HLDG CO LLC
  • US10929019B1 patent drawing
  • US10929019B1 patent drawing
  • US10929019B1 patent drawing

AI summary

Data storage facilities that provide data storage services are typically arranged as either individually accessible drive facilities or multi-carrier drive facilities, but both types must contend with hardware failures that require storage device (e.g., drives) to be replaced. Multi-carrier drive facilities generally have greatly increased drive density, but are confronted with challenges with respect to service operations that are not present for facilities with individually accessible drives. For example, replacing a faulted storage device can entail bringing the faulted storage device as well as other (e.g., non-faulted) storage devices offline during the service operation, which can impact the data storage services. Techniques that improve service for multi-carrier drive facilities are presented. Such techniques can improve coordination between elements that manage the storage facility and those that provide service to faulted storage elements.