SAS Expander Segregating Misbehaving Storage Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Serial Attached SCSI (SAS) systems face challenges in maintaining reliable communication due to higher transmission error rates at faster speeds, leading to delays and inaccessible data, especially between expanders and storage devices, where misbehaving storage devices cause configuration changes and disrupt normal I/O traffic.
Innovation Solution
A method and system that detect errors on links between expanders and storage devices, maintaining error counts, and if above a threshold, isolate the storage device into a segregated zone to prevent further configuration changes and allow normal I/O traffic to resume, enabling autonomous restoration with minimal disruption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error detection and isolation mechanisms are implemented, then system reliability is improved, but device complexity increases
Solution Approach 1:
The SAS domain is segmented into a normal zone and a segregated zone. When a storage device is identified as misbehaving, it is moved from the normal zone to the segregated zone, isolating it from the rest of the system. This segmentation allows the majority of the system to continue operating normally while the problematic device is contained, thus improving reliability without requiring complete system shutdown or complex restructuring.
Solution Approach 2:
The expander acts as an intermediary device that manages the segregation process. It detects errors on links, maintains error counts, and executes commands to move storage devices between zones. The expander mediates between the misbehving storage device and the rest of the system, handling the isolation process automatically and reducing the complexity burden on other system components.
2Productivity
If misbehaving storage devices are isolated quickly, then I/O traffic disruption is minimized, but error detection and processing time increases
Solution Approach 1:
The system performs preliminary error detection by continuously monitoring links and maintaining error counts before complete failures occur. When the error count exceeds a threshold, the isolation process is triggered proactively. This preliminary action allows the system to prepare for and execute isolation quickly, minimizing disruption to I/O traffic while spreading the processing load over time through continuous monitoring.
Solution Approach 2:
The system implements feedback mechanisms where the expander continuously monitors link errors and adjusts its behavior based on error count thresholds. When errors are detected, the system provides feedback by triggering isolation commands, and the status is communicated to initiators and controllers. This feedback loop enables rapid response to misbehaving devices while maintaining system awareness and allowing for dynamic adjustment of isolation decisions.
Data Source
AI summary
A method for maintaining reliable communication on a link between an expander and a storage device is provided. The method includes detecting, by a processor coupled to the link, an error corresponding to the link, and maintaining a count of detected errors for the link, by the processor. The method also includes determining, by the processor, if the count of detected errors is above a first error threshold. If the count of detected errors is not above the first error threshold, then the method repeats the detecting, maintaining, and determining steps. If the count of detected errors is above the first error threshold, then the method provides the processor placing the storage device into a segregated zone.


