Memory Sub-System Capacity Management After Device Failure

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern memory sub-systems face challenges in managing capacity reductions due to device failures, leading to inefficient waste of remaining memory devices and increased costs, as existing fault tolerance mechanisms are limited in handling failures and can affect performance and reliability.

Innovation Solution

The memory sub-system detects failures, communicates with the host system to manage capacity reductions, allowing it to update its configuration to operate at a lower capacity using remaining devices, thereby preserving data and extending the system's operational duration without minimizing internal storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the memory sub-system replaces the entire sub-system when a memory device fails, then reliability is maintained, but cost increases and remaining functional memory devices are wasted

Engineering Contradiction:
Improvesystem reliabilityVSAvoidwaste of remaining memory devices
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The memory sub-system is segmented into multiple independent memory devices, allowing individual devices to be managed separately. When one device fails, only that specific device is deactivated while other devices continue to operate, enabling granular fault isolation and resource utilization

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of discarding the entire memory sub-system upon failure of one device, the system recovers by deactivating only the failed device and redistributing its data to remaining functional devices. This selective discarding approach preserves valuable remaining resources while maintaining system reliability

Inventive Principle:
Principle #34Discarding and recovering

2Reliability

If the memory sub-system deactivates a failed memory device and redistributes data, then capacity is reduced, but the system can continue operating with remaining devices

Engineering Contradiction:
Improvefault toleranceVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The memory sub-system dynamically adjusts its configuration by deactivating failed devices and redistributing data across remaining devices. This dynamic reconfiguration allows the system to adapt to capacity reductions while maintaining operational reliability

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters by modifying the active device configuration and data distribution scheme. When capacity is reduced, the system adjusts its data placement strategy to utilize remaining devices effectively, transforming the static capacity into a dynamically adaptable resource

Inventive Principle:
Principle #35Parameter changes

3Speed

If internal storage space is minimized to maintain performance, then speed is improved, but capacity reduction is exacerbated

Engineering Contradiction:
Improvememory access speedVSAvoidavailable storage capacity
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The system applies different quality levels to different data portions by identifying and preserving critical data that requires fast access while allowing less critical data to be redistributed to remaining capacity. This local quality differentiation maintains performance for essential operations while accommodating overall capacity reduction

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11892909B2Managing capacity reduction due to storage device failure
Publication Date: 2024.02.06 MICRON TECHNOLOGY INC
  • US11892909B2 patent drawing
  • US11892909B2 patent drawing
  • US11892909B2 patent drawing

AI summary

A system and method for managing a reduction in capacity of a memory sub-system. An example method involving a memory sub-system: detecting a failure of at least one memory device of the set, wherein the failure affects stored data; notifying a host system of a change in a capacity of the set of memory devices; receiving from the host system an indication to continue at a reduced capacity; and updating the set of memory devices to change the capacity to the reduced capacity.