RAID SSD Auto-Rebuild and Snapshot Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed storage systems, particularly RAID systems with SSDs, face challenges in maintaining high performance and dependability due to the overhead of manual rebuild operations when an SSD fails, which impacts overall system performance and reliability.

Innovation Solution

The method involves shifting host device functions to SSDs, enabling auto-rebuild and auto-error correction operations without host intervention, and managing snapshot features to maintain data integrity and reduce processing load on the host device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual rebuild operations are performed when an SSD fails, then data integrity is maintained, but system performance and reliability deteriorate due to host device overhead

Engineering Contradiction:
Improvedata integrityVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables SSDs to autonomously perform rebuild operations when failures occur, without requiring host device intervention. Each SSD monitors its own health status and automatically initiates data reconstruction from remaining healthy drives, allowing the system to maintain data integrity while eliminating host overhead during failure recovery events

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements proactive monitoring of SSD health parameters and pre-configures rebuild capabilities before failures occur. By detecting early signs of SSD degradation and having rebuild algorithms ready, the system can quickly restore data integrity without performance degradation, as the rebuild infrastructure is already in place rather than being assembled during critical failure events

Inventive Principle:
Principle #10Preliminary action

2Reliability

If host device performs error detection and correction, then data reliability is improved, but host device processing load increases

Engineering Contradiction:
Improvedata reliabilityVSAvoidhost device processing load
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the error detection and correction functionality into separate modules within each SSD, rather than centralizing these functions in the host device. Each SSD contains its own error monitoring and correction algorithms, allowing distributed processing that maintains data reliability while eliminating the processing burden from the host device

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces SSD-level error management as an intermediary layer between the physical storage medium and the host device. This intermediate error handling capability at the SSD level filters out errors before they reach the host, maintaining data reliability while protecting the host from processing complex error correction tasks

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If snapshot features are managed by host device, then data recovery capability is improved, but system complexity and overhead increase

Engineering Contradiction:
Improvedata recovery capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent enables SSDs to autonomously create and manage snapshots of their own data without host device involvement. Each SSD maintains its own snapshot metadata and can independently restore previous data states, providing robust data recovery capability while eliminating the complexity of centralized snapshot management infrastructure

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11645153B2Method for controlling operations of raid system comprising host device and plurality of SSDs
Publication Date: 2023.05.09 SAMSUNG ELECTRONICS CO LTD
  • US11645153B2 patent drawing
  • US11645153B2 patent drawing
  • US11645153B2 patent drawing

AI summary

Embodiments herein provide a method for controlling operations of a Redundant Array of Independent Disks (RAID) data storage system comprising a host device and a plurality of solid-state drives (SSDs). The method includes performing, by the at least one SSD, recovery of lost data by performing the auto-rebuild operation. The method also includes performing by the at least one SSD, the auto-error correction operation based on the IO error. The method also includes creating a snapshot of an address mapping table by all SSDs of the plurality of SSDs in the RAID data storage system. The auto-rebuild operation, the auto-error correction operation and the creation the snapshot of the address mapping table are all performed without the intervention from the host device.