Active-standby Storage Controllers for Cross-site Data Redundancy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In cloud storage systems, communication delays between geographically separated availability zones lead to deteriorated I/O performance and increased costs due to high communication volumes.

Innovation Solution

A distributed storage system with redundancy groups across sites, where an active storage controller processes data locally and stores redundant data at another site, allowing a standby controller to take over in case of failure, thus maintaining data locality and reducing cross-site communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is stored across multiple geographically separated availability zones for high availability, then system reliability is improved, but communication delay increases and I/O performance deteriorates

Engineering Contradiction:
Improvehigh availabilityVSAvoidI/O performance
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The system segments storage controllers into active and standby roles within redundancy groups, and segments data into primary and redundant portions. The active storage controller processes I/O requests locally without needing to communicate with remote availability zones, while redundant data is asynchronously replicated. This segmentation allows local operations to proceed at full speed while maintaining cross-zone redundancy for failover capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary replication of redundant data to remote availability zones before failures occur. The active storage controller asynchronously replicates data to standby storage controllers in other availability zones in advance, so that when a failure occurs, the standby controller can immediately take over without requiring real-time communication across zones, thus maintaining I/O performance while ensuring high availability.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If redundant data is replicated across multiple sites for disaster recovery, then system reliability is improved, but communication volume increases and costs increase

Engineering Contradiction:
Improvedisaster recovery capabilityVSAvoidcommunication volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system performs preliminary replication of redundant data to remote availability zones before failures occur. The active storage controller asynchronously replicates data to standby storage controllers in other availability zones in advance, so that when a failure occurs, the standby controller can immediately take over without requiring real-time communication across zones, thus maintaining I/O performance while ensuring high availability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates redundant copies of data in standby storage controllers located in different availability zones. Instead of maintaining continuous real-time synchronization that would generate high communication volume, the system uses asynchronous replication where the standby controller receives and stores copies of data independently, reducing the need for frequent cross-zone communication while ensuring disaster recovery capability.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240220378A1Information processing system and information processing method
Publication Date: 2024.07.04 HITACHI VANTARA LTD
  • US20240220378A1 patent drawing
  • US20240220378A1 patent drawing
  • US20240220378A1 patent drawing

AI summary

Proposed are a highly available information processing system and information processing method capable of withstanding a failure in units of sites. A redundancy group including a plurality of the storage controllers installed in different sites is formed, and the redundancy group includes an active state storage controller which processes data, and a standby state storage controller which takes over processing of the data if a failure occurs in the active state storage controller, and the active state storage controller executes processing of storing the data from a host application installed in the same site in the storage device installed in that site, and storing redundant data for restoring data stored in a storage device of a same site in the storage device installed in another site where a standby state storage controller of a same redundancy group is installed.