Storage Sled Component Identification for Targeted Replacement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face inefficiencies due to the practice of discarding all storage devices on a sled when one fails, leading to waste and resource-intensive failure analysis, as all devices are treated as faulty despite only one being defective, and transient failures can go undetected.
Innovation Solution
Implementing a storage system with field replaceable units (FRUs) that include memory for identifying and storing information about failed components, allowing for targeted replacement and maintenance, where the FRU with the faulty device can be isolated and serviced while operational devices are reused.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all storage devices on a sled are treated as faulty when one fails, then the storage system can be quickly restored by replacing the entire sled, but operational storage devices are wasted and resource consumption increases
Solution Approach 1:
The invention segments the failure identification to the individual storage device level rather than treating the entire sled as faulty. Each storage device has its own identifier and failure status can be independently determined, allowing only the failed device to be replaced while preserving other operational devices on the same sled.
Solution Approach 2:
The invention introduces an intermediary identification mechanism (identifier stored in memory associated with the storage device or sled) that enables precise tracking of which specific device failed. This intermediary allows the system to distinguish between failed and operational devices, preventing unnecessary disposal of functional storage devices.
2Measurement precision
If full failure analysis is performed on each storage device, then accurate identification of faulty devices is possible, but manpower and resource consumption increase significantly
Solution Approach 1:
The invention performs preliminary action by pre-storing identifiers and failure information in memory associated with each storage device or sled before failure occurs. When a failure happens, the system can quickly retrieve pre-stored identification information rather than performing time-consuming analysis, thus maintaining high accuracy while improving efficiency.
Solution Approach 2:
The storage devices or sleds serve themselves by maintaining their own identifiers and failure information in locally associated memory. This self-service mechanism eliminates the need for external manual analysis, allowing the system to automatically identify failed devices without consuming significant manpower or resources.
3Ease of repair
If storage devices are mounted on replaceable sleds, then serviceability and packaging are facilitated, but all devices on a sled must be replaced together when one fails
Solution Approach 1:
The invention segments the replacement unit from the sled level down to the individual storage device level. While devices remain mounted on sleds for ease of handling, the identification and replacement process is segmented to target only the specific failed device, allowing operational devices on the same sled to be preserved and reused.
Solution Approach 2:
The invention applies local quality by associating failure identification information specifically with the failed storage device rather than the entire sled. This allows differential treatment of devices on the same sled - the failed device is replaced while operational devices are retained, optimizing resource utilization while maintaining serviceability.
Data Source
AI summary
The present invention is directed to a data storage system utilizing a number of data storage devices. The data storage system features one or more storage device sleds, which may each carry multiple storage devices. Each storage device sled and its interconnected storage devices may comprise a field replaceable unit. In response to the detection of a failure associated with a field replaceable unit, information related to that failure may be stored in memory or storage associated with the field replaceable unit. Repair personnel may access the stored information in order to positively identify the failed component of the field replaceable unit in connection with the repair or replacement of that component, in order to return the field replaceable unit to service.


