Memory Power Fault Resilience via Controller Mirroring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Information handling systems face downtime and increased maintenance costs due to automatic shutdowns triggered by power faults in memory units, which can be caused by transient issues or failures, leading to unnecessary system reboots and requiring technician intervention.
Innovation Solution
A controller, such as a complex programmable logic device (CPLD), monitors memory units for power faults and intervenes to prevent automatic shutdowns by using memory mirroring to switch to healthy memory units, and only arms power fault detection mechanisms when a stable power good signal is detected, thereby avoiding false triggers and allowing provisional reactivation of faulty memories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If automatic shutdown is triggered by power fault detection in memory units, then system reliability is improved by preventing damage from power failures, but system downtime increases and requires technician intervention
Solution Approach 1:
The system performs preliminary actions by detecting power faults and switching to backup memory units before the power failure causes system damage or complete shutdown. The controller monitors power good signals and proactively switches to mirrored memory units when voltage anomalies are detected, preventing catastrophic failure and avoiding technician intervention.
Solution Approach 2:
The system implements beforehand cushioning through memory mirroring redundancy. Healthy mirrored memory units are prepared in advance as backups, so when a power fault occurs in the primary memory, the system can immediately switch to the pre-prepared backup without shutdown. This cushioning mechanism absorbs the shock of power failures and maintains continuous operation.
2Reliability
If power fault detection is continuously monitored, then system reliability is improved by detecting actual power failures, but false triggers from transient issues cause unnecessary shutdowns
Solution Approach 1:
The system applies preliminary anti-action by implementing a delay mechanism that prevents immediate shutdown responses to transient power faults. When a power fault is detected, the system waits for a predetermined period to see if the fault persists before triggering a shutdown. This preliminary anti-action counteracts false triggers from temporary voltage fluctuations while maintaining detection of genuine power failures.
Solution Approach 2:
The detection mechanism dynamically adjusts its response based on the persistence of power faults. The system transitions from immediate shutdown response to delayed response when faults are transient, and to actual shutdown only when faults persist beyond the predetermined period. This dynamic behavior optimizes the balance between detecting real failures and avoiding false positives.
Data Source
AI summary
A controller of an information handling system may detect a power fault event for one or more of a plurality of memories configured to operate in a memory mirroring mode. The controller may deactivate the one or more of the plurality of memories by mapping the one or more of the plurality of memories out from usage without rebooting the information handling system based, at least in part, on the detection of the power fault event and the received notification.


