Disk Array Expander Controller Redundancy for Failure Isolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional disk array systems face issues where a blocked drive box can lead to all subsequent drive boxes being blocked, resulting in data unavailability and performance deterioration during RAID group recovery, with wide spread failure tolerance and recovery process impacts.
Innovation Solution
A disk array system with redundant controllers and daisy chain-connected chassis, where expander controllers from non-blocked drive boxes can access and reroute data through alternative paths, reducing the failure range and maintaining system availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If drive boxes are connected in daisy chain mode to expand the storage system, then the system capacity and connectivity are improved, but a failure in one drive box can block all subsequently arranged drive boxes
Solution Approach 1:
The system segments the connection paths by providing multiple independent routes from the controller to drive boxes. Instead of a single daisy-chain path, the controller can reach drive boxes through different expander controllers, creating isolated segments that prevent failure propagation across the entire system.
Solution Approach 2:
The patent adds a dimensional aspect to the connection topology by introducing multiple expander controllers that provide alternative paths. This transforms the linear daisy-chain structure into a multi-dimensional network where drive boxes can be accessed through different routes, effectively adding redundancy in the connection dimension.
2Reliability
If data recovery is performed based on RAID group configuration when drive box failure occurs, then data can be recovered, but the recovery processing must be performed in all subsequently arranged drive boxes which deteriorates performance
Solution Approach 1:
The failure impact is segmented and localized to only the affected drive box and its directly connected expander controller. Other drive boxes accessed through alternative paths remain operational and do not require recovery processing, thus improving overall system performance during failure recovery scenarios.
Solution Approach 2:
Instead of performing recovery processing in all subsequently arranged drive boxes, the system applies recovery processing only to the minimal necessary scope - the specific drive box and expander controller affected by the failure. This partial action approach maintains data recovery capability while avoiding unnecessary performance deterioration.
3Difficulty of detecting and measuring
If a drive box is blocked due to failure, then the failure section can be identified, but all subsequently arranged drive boxes connected to the blocked drive box will also be blocked
Solution Approach 1:
The connection topology is segmented into multiple independent paths, allowing the system to identify and isolate the specific failed component without affecting other segments. When a drive box or expander controller fails, only the specific path through that component is blocked, while alternative paths remain open for accessing other drive boxes.
Solution Approach 2:
Multiple expander controllers act as intermediaries that provide alternative routing paths. When one expander controller or its connected drive box fails, data can still be accessed by routing through other expander controllers, preventing complete data unavailability and allowing continued system operation.
Data Source
AI summary
According to the conventional art disk array system, when a drive box of a first chassis is blocked, all the drive boxes of a second chassis connected subsequently therefrom will be blocked, and the data in the drive boxes arranged subsequently therefrom cannot be recovered based on the RAID configuration. Even if the data could be recovered based on RAID configuration, it is necessary to perform the recovery process based on RAID in all the subsequently arranged drive boxes, according to which the performance is deteriorated. The present system stores a first drive box in a first chassis out of a plurality of chassis, and a second drive box and a third drive box are stored in a second chassis. One of a plurality of expander controllers within the first drive box is connected to an expander controller in the second drive box, and the other expander controller is connected to an expander controller in the third drive box.


