Memory Power Fault Resilience via Controller Mirroring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Information handling systems face downtime and increased maintenance costs due to automatic shutdowns triggered by power faults in memory units, which can be caused by transient issues or failures, leading to unnecessary system reboots and requiring technician intervention.

Innovation Solution

A controller, such as a complex programmable logic device (CPLD), monitors memory units for power faults and intervenes to prevent automatic shutdowns by using memory mirroring to switch to healthy memory units, and only arms power fault detection mechanisms when a stable power good signal is detected, thereby avoiding false triggers and allowing provisional reactivation of faulty memories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If automatic shutdown is triggered by power fault detection in memory units, then system reliability is improved by preventing damage from power failures, but system downtime increases and requires technician intervention

Engineering Contradiction:
Improvesystem reliabilityVSAvoidsystem downtime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by detecting power faults and switching to backup memory units before the power failure causes system damage or complete shutdown. The controller monitors power good signals and proactively switches to mirrored memory units when voltage anomalies are detected, preventing catastrophic failure and avoiding technician intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements beforehand cushioning through memory mirroring redundancy. Healthy mirrored memory units are prepared in advance as backups, so when a power fault occurs in the primary memory, the system can immediately switch to the pre-prepared backup without shutdown. This cushioning mechanism absorbs the shock of power failures and maintains continuous operation.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

2Reliability

If power fault detection is continuously monitored, then system reliability is improved by detecting actual power failures, but false triggers from transient issues cause unnecessary shutdowns

Engineering Contradiction:
Improvefault detection accuracyVSAvoiddetection mechanism complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system applies preliminary anti-action by implementing a delay mechanism that prevents immediate shutdown responses to transient power faults. When a power fault is detected, the system waits for a predetermined period to see if the fault persists before triggering a shutdown. This preliminary anti-action counteracts false triggers from temporary voltage fluctuations while maintaining detection of genuine power failures.

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The detection mechanism dynamically adjusts its response based on the persistence of power faults. The system transitions from immediate shutdown response to delayed response when faults are transient, and to actual shutdown only when faults persist beyond the predetermined period. This dynamic behavior optimizes the balance between detecting real failures and avoiding false positives.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11742054B2Memory power fault resilience in information handling systems
Publication Date: 2023.08.29 DELL PROD LP
  • US11742054B2 patent drawing
  • US11742054B2 patent drawing
  • US11742054B2 patent drawing

AI summary

A controller of an information handling system may detect a power fault event for one or more of a plurality of memories configured to operate in a memory mirroring mode. The controller may deactivate the one or more of the plurality of memories by mapping the one or more of the plurality of memories out from usage without rebooting the information handling system based, at least in part, on the detection of the power fault event and the received notification.