Two-Node Storage System Automatic Data Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current two-node storage systems face challenges in maintaining high availability due to the need for manual intervention and prolonged data recovery times when a storage device fails, especially in configurations using RAID1 between nodes, which limits scalability and efficiency.

Innovation Solution

Implementing a method that creates mirrored logic unit groups across two storage nodes, allowing for automatic data recovery from the second node when a failure occurs, thereby reducing recovery time and eliminating the need for human intervention by distributing storage slices across multiple devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If RAID1 is used between storage nodes for high availability, then data redundancy is improved, but data recovery time increases and manual intervention is required

Engineering Contradiction:
Improvedata redundancyVSAvoiddata recovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-establishing mirrored logic unit groups across storage nodes before failures occur. When a storage device fails, the system can immediately activate the pre-configured mirrored data from another node, eliminating the need for time-consuming manual data reconstruction and enabling automatic failover.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service through automatic failure detection and data recovery mechanisms. When a storage device fails, the system automatically detects the failure, retrieves mirrored data from the corresponding logic unit group on another storage node, and restores data without requiring manual intervention, thereby reducing recovery time while maintaining redundancy.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If storage slices are distributed across multiple devices, then system scalability is improved, but system complexity increases

Engineering Contradiction:
Improvesystem scalabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing storage devices into multiple storage slices and distributing them across different storage nodes. Each logic unit group spans multiple storage slices across different nodes, allowing the system to scale by adding more nodes while maintaining data redundancy and simplifying the management of distributed storage through logical grouping.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If manual intervention is required for data recovery, then data accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvedata accuracyVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system implements self-service through automatic failure detection and data recovery mechanisms. When a storage device fails, the system automatically detects the failure, retrieves mirrored data from the corresponding logic unit group on another storage node, and restores data without requiring manual intervention, thereby reducing recovery time while maintaining redundancy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11269745B2Two-node high availability storage system
Publication Date: 2022.03.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11269745B2 patent drawing
  • US11269745B2 patent drawing
  • US11269745B2 patent drawing

AI summary

Aspects of the present invention disclose a method for a two-node storage system. The method includes one or more processors creating a plurality of first logic unit groups in a first storage node of a storage system. The method further includes mapping each of the plurality of first logic unit groups to a number of storage slices from different storage devices in the first storage node. The method further creating a plurality of second logic unit groups in a second storage node of the storage system, by mirroring storage slices from a storage device in the first storage node to multiple storage devices in the second storage node. In response to identifying a failure of a first storage device in the first storage node, the method further includes recovering lost data based on data in the second storage node.