Application-Managed Fault Detection for Cross-Region Object Stores

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional storage systems face challenges in maintaining high availability and disaster recovery for cross-region replicated object stores, particularly during communications outages, where existing solutions often fail to ensure seamless data replication and fault handling across regions.

Innovation Solution

The implementation of application-managed fault handling and replication methods, which include specific flow charts and infrastructure management techniques to control cross-region replicated object stores, ensuring data integrity and availability through proactive data rebuilding and redundancy mechanisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional storage systems use automatic replication mechanisms for cross-region data storage, then data redundancy is improved, but system complexity and difficulty in handling communication outages increase

Engineering Contradiction:
Improvedata availabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables self-service through automated fault detection and application-managed fault handling. The primary object store automatically detects faults and notifies the application, which then manages the fault handling process including initiating data rebuilding to the secondary object store, eliminating the need for complex manual intervention mechanisms

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system segments fault handling responsibilities between the storage system (automatic fault detection and notification) and the application layer (fault handling and data rebuilding). This segmentation simplifies each component's functionality while maintaining overall system reliability

Inventive Principle:
Principle #1Segmentation

2Reliability

If traditional storage systems implement cross-region replication, then disaster recovery capability is improved, but performance degradation during communications outages occurs

Engineering Contradiction:
Improvedisaster recovery capabilityVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary actions by pre-configuring secondary object stores in different geographic regions and pre-establishing replication capabilities. When communications outages occur, the system can immediately initiate data rebuilding to the secondary location without performance degradation, as the infrastructure is already in place

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary notification mechanism where the primary object store notifies the application of faults, which then coordinates the data rebuilding process. This intermediary approach allows the system to handle communications outages gracefully by separating the detection and execution phases of fault handling

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If manual fault handling procedures are used in cross-region replicated object stores, then data integrity can be maintained, but response time and operational efficiency deteriorate

Engineering Contradiction:
Improvedata integrityVSAvoidfault response time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements feedback mechanisms where the primary object store automatically detects faults and notifies the application, which then initiates data rebuilding to the secondary object store. This automated feedback loop ensures rapid response to faults while maintaining data integrity through application-managed fault handling procedures

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary setup by configuring secondary object stores and replication capabilities in advance. When faults occur, the pre-configured system can immediately execute data rebuilding operations without manual intervention, reducing fault response time while maintaining data integrity

Inventive Principle:
Principle #10Preliminary action

4Reliability

If seamless data replication is implemented across regions, then high availability is improved, but system complexity and difficulty in managing communications outages increase

Engineering Contradiction:
Improvehigh availabilityVSAvoidreplication management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system achieves self-service through automated fault detection by the primary object store and automatic notification to the application, which then manages the replication process. This eliminates the need for complex manual management of cross-region replication while maintaining high availability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system segments replication management into automated components (fault detection and notification) and application-managed components (fault handling and data rebuilding). This segmentation simplifies the overall system complexity while maintaining seamless data replication across regions

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230393927A1Application-Managed Fault Detection For Cross-Region Replicated Object Stores
Publication Date: 2023.12.07 PURE STORAGE INC
  • US20230393927A1 patent drawing
  • US20230393927A1 patent drawing
  • US20230393927A1 patent drawing

AI summary

Application-managed fault detection for cross-region replicated object stores is disclosed. An embodiment includes determining, by a first storage system among a plurality of storage systems replicating an object store, a faulted state in response to identifying a fault that prevents replication of updates to the object store to at least a second storage system of the plurality of storage systems; providing, through an API, an indication that the first storage system has entered the faulted state; and receiving a request indicating how to proceed in the presence of the fault.