Storage Array Link Error Detection and Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Detecting and localizing congestion and failures in complex data center systems, such as storage area networks, is a difficult manual process due to their intricate nature and multi-layered network interconnections.

Innovation Solution

A method and apparatus that monitor IO failures and link errors at storage array ports, initiator ports of host servers, and adjacent switch ports, enabling remote detection and localization of issues through ongoing monitoring and mapping, with a congestion and failure detector and localizer running on compute nodes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Difficulty of detecting and measuring

If manual detection methods are used for congestion and failures in data center systems, then system complexity is managed, but detection speed and accuracy deteriorate

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The system segments the detection task by monitoring at specific points (storage array ports, host server initiator ports, and adjacent switch ports) rather than attempting to monitor the entire complex system at once. This segmentation allows accurate detection of link errors and IO failures at critical locations while managing overall system complexity through targeted monitoring.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary monitoring system that collects data from multiple sources (storage array, host servers, and switches) and processes this information to detect and localize problems. This intermediary layer simplifies the detection process by centralizing monitoring functions and providing a unified approach to analyzing complex system behavior.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If comprehensive monitoring of all links is implemented, then detection precision improves, but monitoring complexity and resource consumption increase

Engineering Contradiction:
Improvelocalization precisionVSAvoidmonitoring complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies local quality by focusing monitoring resources on specific critical points where link errors and IO failures are most likely to occur - namely storage array ports, host server initiator ports, and adjacent switch ports. This targeted approach achieves precise localization of problems without the complexity of monitoring every possible link in the system.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent adds a dimensional approach by monitoring not just the physical links but also the logical paths and error patterns across multiple layers (storage array, host server, switch). This multi-dimensional monitoring enables precise localization of congestion and failures by analyzing data from multiple perspectives simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of time

If real-time monitoring of IO failures and link errors is implemented, then response time improves, but system overhead increases

Engineering Contradiction:
Improveresponse timeVSAvoidmonitoring overhead
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The system implements periodic monitoring of IO failures and link errors at key points rather than continuous exhaustive monitoring of all system components. This periodic action at critical locations achieves timely detection and response while reducing overall monitoring overhead by focusing resources on when and where errors are most likely to occur.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS11809268B1Discovering host-switch link and ISL issues from the storage array
Publication Date: 2023.11.07 DELL PROD LP
  • US11809268B1 patent drawing
  • US11809268B1 patent drawing
  • US11809268B1 patent drawing

AI summary

A congestion and failure detector and localizer running on a storage array locally monitors ports of the storage array for IO failures and local link errors and remotely monitors ports of host initiators and host-side and storage array-side switches for link errors. Based on whether the local link error rate is increasing at any ports, whether IO failures are associated with a single host initiator port, and whether link error rate is increasing on both the host initiator and initiator-side switch, the congestion and failure detector and localizer generates a flag indicating either a physical link problem between the storage array and adjacent switch, ISL physical issue or spreading congestion, host initiator-side physical link problem, or path congestion. The flag identifies the storage array port, issue type, link, fabric name, and host initiator.