Automated Stall Detection in Dispersed Storage Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems lack an efficient and effective way to handle unresponsive or semi-responsive storage devices in distributed storage networks, which can lead to data loss and system instability due to conditions like kernel hangs, deadlocks, and hardware errors.

Innovation Solution

A dispersed storage network with a managing unit and integrity processing unit that performs error encoding and decoding, and automated detection and recovery of stalled processes, using Cauchy Reed-Solomon encoding and external monitoring to identify and correct failing storage units, ensuring data integrity and availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional storage systems are used without automated monitoring, then device complexity is reduced, but reliability deteriorates due to undetected storage unit failures

Engineering Contradiction:
Improvestorage system reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The storage system performs self-diagnosis through automated stall detection where storage units monitor each other's responsiveness. The system automatically detects failed storage units and triggers recovery processes without external intervention, enabling the system to self-heal and maintain reliability while minimizing manual complexity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous feedback loops where storage units send responsiveness indicators to the dispersed storage network. This feedback mechanism allows the system to detect stalled processes in real-time and automatically initiate recovery actions, improving reliability through proactive monitoring while keeping the complexity manageable through automated responses

Inventive Principle:
Principle #23Feedback

2Reliability

If automated monitoring and recovery mechanisms are implemented, then reliability improves through detection of stalled processes, but device complexity increases due to additional monitoring and recovery components

Engineering Contradiction:
Improvestorage unit responsivenessVSAvoidmonitoring and recovery system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system introduces intermediary components including a stall detector module and recovery orchestrator that mediate between storage units and the dispersed storage network. These intermediaries simplify the overall system architecture by centralizing monitoring and recovery logic, improving reliability through specialized components while managing complexity through clear separation of concerns

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The monitoring and recovery system is segmented into distinct functional modules: stall detection, error indication, recovery determination, and corrective action execution. This segmentation allows each component to be independently optimized and maintained, improving reliability through specialized functionality while reducing overall system complexity through modular design

Inventive Principle:
Principle #1Segmentation

3Reliability

If error correction encoding is used for all data, then data integrity is improved, but use of energy increases due to additional processing and storage requirements

Engineering Contradiction:
Improvedata integrityVSAvoidprocessing energy
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system applies error correction encoding selectively rather than universally. Cauchy Reed-Solomon encoding is applied to critical data segments that require enhanced protection, while less critical data uses standard encoding. This partial application maintains data integrity for essential information while reducing the overall energy burden of encoding and decoding operations across the entire storage system

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10698778B2Automated stalled process detection and recovery
Publication Date: 2020.06.30 PURE STORAGE INC
  • US10698778B2 patent drawing
  • US10698778B2 patent drawing
  • US10698778B2 patent drawing

AI summary

A dispersed storage network (DSN) includes multiple storage units. A processing unit included in the DSN issues an access request to one of the storage units, and identifies the storage unit as a failing storage unit based, at least in part, on a rate of growth of a network queue associated with the storage unit. the processing unit then issues an error indicator to a recovery unit for further action.