BIOS Error Tracking for NV DIMM Stability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Systems with non-volatile dual in-line memory modules (NV DIMMs) face instability due to continuous system crashes caused by reinstating bad data during power restoration, leading to cycles of crashes and restarts until faulty NV DIMMs are removed.
Innovation Solution
A BIOS chip tracks error counts for memory modules and executes corrective actions, such as reinitialization or disabling, based on predefined thresholds to prevent bad data from being reinstated, ensuring system stability by managing memory errors proactively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Duration of action of stationary object
If data is copied from volatile to non-volatile memory components in NV DIMM, then data persistence is improved, but system stability deteriorates due to reinstating bad data after power restoration
Solution Approach 1:
The patent applies preliminary action by checking memory modules for errors and storing error information in non-volatile memory before power loss occurs. During power restoration, the BIOS reads this pre-stored error information and takes corrective actions (such as clearing error flags or reinitializing memory) before bad data is reinstated, thereby preventing system crashes while maintaining data persistence capability
Solution Approach 2:
The patent implements feedback by continuously monitoring memory module status and storing error information in non-volatile memory. When power is restored, the system reads the stored error information and adjusts its behavior accordingly - if errors are detected, corrective actions are taken before data reinstatement. This closed-loop feedback mechanism prevents bad data from causing system instability while preserving the benefits of non-volatile data persistence
2Reliability
If error data is tracked and corrective actions are taken based on error counts, then system stability is improved, but device complexity increases due to additional tracking and decision-making mechanisms
Solution Approach 1:
The patent applies self-service by enabling the memory system to automatically monitor its own health status, store error information in non-volatile memory, and perform self-correction during power restoration. The BIOS autonomously reads error flags and executes appropriate corrective actions without requiring external intervention or complex external control systems, thereby improving reliability while minimizing additional complexity
Solution Approach 2:
The patent merges the error tracking function with the existing non-volatile memory component of the NV DIMM. Instead of adding separate complex tracking systems, the error information is stored in the same non-volatile memory that already exists for data persistence. This consolidation achieves error monitoring and corrective capabilities while reusing existing hardware resources, thus improving reliability without proportionally increasing device complexity
Data Source
AI summary
In one example in accordance with the present disclosure, a system for handling memory errors includes a memory module having volatile components and non-volatile components. The system includes a BIOS chip having BIOS code and a BIOS non-volatile (NV) memory. The BIOS NV memory stores error data associated with the memory module that was stored prior to a power-on or reset of the system. The system includes a processor to execute the BIOS code to, after the power-on or reset of the system end before an operating system is loaded; (1) read, from the BIOS NV memory, the error data; and (2) determine, based on the error data, whether to take a corrective action with respect to the memory module.


