Distributed Storage Volume Recovery During Site Failure

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed storage systems face challenges in automatically creating and maintaining distributed storage volumes during site failures, array failures, or inter-site network failures, which requires manual administrator intervention and is error-prone.

Innovation Solution

A method and system that automatically detect array unavailability, create a local part of the distributed volume, export it to an available array, continue processing I/O requests, and upon array reavailability, create a remote part and perform consistency processing to ensure compliance with distributed storage requirements, using a RAID system to maintain data consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If distributed storage volumes are not automatically created during site failure, array failure or inter-site network failure, then manual administrator intervention is required which is error-prone, but the system cannot maintain automated operation and data availability

Engineering Contradiction:
Improveautomated volume creationVSAvoiderror-prone manual process
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The system automatically detects array unavailability and creates local parts of distributed volumes without administrator intervention. The storage system self-manages the entire workflow including detecting failures, creating local volume parts, exporting to available arrays, and performing consistency processing when arrays return.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by automatically creating local parts of distributed volumes on available arrays before the failed arrays return. This preliminary creation allows I/O processing to continue immediately rather than waiting for manual intervention or array recovery.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If manual workflow steps are required to create volumes on available arrays during failures, then the process is complex and error-prone, but automation is not implemented

Engineering Contradiction:
Improvemanual workflow complexityVSAvoidautomated processing
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

Multiple manual operations are merged into a single automated process. The system combines failure detection, volume creation, data export, I/O processing continuation, and consistency management into one unified automated workflow that executes without administrator intervention.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The complex manual workflow steps are extracted from the administrator's responsibilities and transferred to the storage system's automated processes. The system independently handles detecting array unavailability, creating local volume parts, managing exports, and performing consistency processing.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If distributed storage processing stops during array unavailability, then data consistency is maintained, but I/O requests cannot be processed

Engineering Contradiction:
Improvedata consistencyVSAvoidI/O processing continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The distributed volume is segmented into local and remote parts. When remote arrays are unavailable, the system creates and uses local volume parts that can independently process I/O requests. This segmentation allows continued operation with local data while maintaining the ability to restore full distributed functionality when arrays return.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Local volume parts act as intermediaries between I/O requests and the unavailable remote arrays. These local parts enable continued I/O processing by providing local storage access, while consistency processing later reconciles any differences when remote arrays become available again.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If consistency processing is not performed automatically when arrays return, then compliance with distributed storage requirements may be violated, but manual intervention is required

Engineering Contradiction:
Improvedistributed storage complianceVSAvoidautomatic consistency processing
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The system implements feedback by automatically detecting when previously unavailable arrays return to availability. Upon detecting array reavailability, the system triggers consistency processing to compare and reconcile data between local and remote volume parts, ensuring compliance with distributed storage requirements.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary consistency processing automatically when arrays return, before requiring any manual intervention. This ensures that data compliance is verified and restored immediately upon array recovery, maintaining the integrity of the distributed storage system.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10970181B2Creating distributed storage during partitions
Publication Date: 2021.04.06 EMC IP HLDG CO LLC
  • US10970181B2 patent drawing
  • US10970181B2 patent drawing
  • US10970181B2 patent drawing

AI summary

A system and method are provided for processing to create distributed volume in a distributed storage system during a failure that has partitioned the distributed volume (e.g. an array failure, a site failure and/or an inter-site network failure). In an embodiment, the system described herein may provide for continuing distributed storage processing in response to I/O requests from a source by creating the local parts of the distributed storage during the failure, and, when the remote site or inter-site network return to availability, the remaining part of the distributed volume is automatically created. The system may include an automatic rebuild to make sure that all parts of the distributed volume are consistent again. The processing may be transparent to the source of the I/O requests.