Storage Access Path Management for Disk Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional disk array access techniques fail to effectively handle multiple disk failures during powered operations, leading to poor user experience and potential storage device crashes due to congestion from unfinished data writes.

Innovation Solution

A method that determines if an access timeout is caused by a physical connection failure, suspends access for a scheduled time period to prevent further data influx, and resumes access via a secondary path, thereby protecting the storage device from crashing and improving stability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional disk array access techniques are used, then data storage capacity is improved, but system reliability deteriorates when multiple disk failures occur during powered operations

Engineering Contradiction:
Improvedata storage capacityVSAvoidsystem reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system performs preliminary actions by detecting disk failures and suspending access to affected paths before congestion can occur. The controller proactively identifies failed disks and prevents further write operations to damaged paths, stopping the harmful process before it causes system-wide congestion and crash.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The controller acts as an intermediary between the host and storage devices, managing access paths and detecting failures. It mediates by intercepting access requests, identifying failed paths, and routing operations through healthy paths while suspending access to damaged disks, preventing congestion from propagating through the entire array.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If access continues without suspension after timeout, then productivity is maintained, but harmful effects increase due to congestion from unfinished data writes

Engineering Contradiction:
Improveaccess continuityVSAvoidcongestion from unfinished data writes
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The system applies preliminary anti-action by suspending access to failed paths before congestion can develop. When a timeout indicates disk failure, the controller immediately stops write operations to that path, preventing the accumulation of unfinished data writes that would otherwise cause congestion and potential system crash.

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The harmful element of unfinished data writes is extracted by suspending access to failed paths. The controller identifies and isolates the problematic paths by stopping access operations, removing the source of congestion from the system while allowing healthy paths to continue operating.

Inventive Principle:
Principle #2Taking out (Extraction)

3Stability of the object's composition

If access is suspended for a scheduled time period, then stability is improved by preventing crashes, but loss of time occurs during the suspension period

Engineering Contradiction:
Improveaccess stabilityVSAvoidsuspension time period
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The system implements periodic action by suspending access for a scheduled time period and then resuming operations. This periodic suspension allows the system to stabilize after detecting disk failures, preventing crashes while ultimately restoring access to maintain productivity. The temporary interruption is followed by resumption of normal operations.

Inventive Principle:
Principle #19Periodic action

4Reliability

If multiple paths are used for redundancy, then reliability is improved, but device complexity increases due to path management requirements

Engineering Contradiction:
Improveaccess reliabilityVSAvoidpath management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system applies self-service by enabling automatic failure detection and path selection without external intervention. The controller autonomously monitors access paths, detects timeouts indicating disk failures, suspends access to failed paths, and resumes operations when appropriate, eliminating the need for complex manual path management while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230273856A1Method, electronic device, and computer program product for accessing storage device
Publication Date: 2023.08.31 DELL PROD LP
  • US20230273856A1 patent drawing
  • US20230273856A1 patent drawing
  • US20230273856A1 patent drawing

AI summary

Embodiments of the present disclosure provide a method, an electronic device, and a computer program product that involve accessing a storage device. The method includes determining, in response to an access to the storage device via a first path being determined as timeout, whether an error on the first path is of a first error type. The method further includes causing the access to be suspended for at least a scheduled time period if the error on the first path is of the first error type. The method further includes resuming the access via a second path after the scheduled time period expires. With embodiments of the present disclosure, the capability of processing the problem of failure in accessing a storage device and the stability in accessing a storage device can be improved.