Multi-drive sled servicing via lock-step isolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multi-carrier drive facilities face challenges in achieving high drive density while minimizing the impact of service operations, as retracting a sled to access a failed drive decouples all storage devices on that sled, including non-faulty ones, often requiring the entire storage array to be brought offline.
Innovation Solution
Implementing a lock-step process between the operating system and hardware elements, with a user interface that indicates when it is safe to begin service, allowing for selective decoupling of faulty drives and enabling preparations for service without affecting operational drives, ensuring minimal disruption to the storage facility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of repair
If a sled is retracted to access a failed drive in a multi-carrier drive facility, then service accessibility is improved, but all storage devices on that sled are decoupled from the storage array including non-faulty ones, reducing productivity
Solution Approach 1:
The system segments the decoupling operation to affect only the specific faulty drive rather than the entire sled. The lock-step process enables individual drive isolation while keeping other drives on the same sled connected to the storage array, thus maintaining productivity while enabling service access.
Solution Approach 2:
The system performs preliminary actions by establishing a lock-step relationship between the operating system and hardware elements before service begins. This preliminary coordination allows the system to prepare for selective decoupling, ensuring that only necessary components are disconnected while others remain operational.
2Reliability
If the entire storage array is brought offline for service operations, then data integrity is protected, but service time and productivity are reduced
Solution Approach 1:
The system applies local quality by implementing service isolation at the drive level rather than array level. The lock-step mechanism enables different states for different drives - the faulty drive is decoupled for service while other drives remain coupled and operational, thus protecting data integrity locally without requiring global offline status.
Solution Approach 2:
Instead of bringing down the entire storage array (excessive action), the system performs partial action by isolating only the specific faulty drive. This partial decoupling approach maintains sufficient data integrity protection while minimizing service time and productivity loss.
3Quantity of substance
If multiple storage devices are housed on a single sled to increase drive density, then space efficiency is improved, but service complexity increases as all devices must be accessed together
Solution Approach 1:
The system segments the service control mechanism to operate at the individual drive level within the sled architecture. The lock-step process creates independent control paths that allow service personnel to address specific drives without affecting others on the same sled, thus maintaining high drive density while reducing service complexity.
Data Source
AI summary
Data storage facilities that provide data storage services are typically arranged as either individually accessible drive facilities or multi-carrier drive facilities, but both types must contend with hardware failures that require storage device (e.g., drives) to be replaced. Multi-carrier drive facilities generally have greatly increased drive density, but are confronted with challenges with respect to service operations that are not present for facilities with individually accessible drives. For example, replacing a faulted storage device can entail bringing the faulted storage device as well as other (e.g., non-faulted) storage devices offline during the service operation, which can impact the data storage services. Techniques that improve service for multi-carrier drive facilities are presented. Such techniques can improve coordination between elements that manage the storage facility and those that provide service to faulted storage elements.


