Memory Error Data Preservation via Non-Volatile Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems face challenges in accurately logging and handling uncorrectable hardware errors in memory systems accessed by microprocessors, as critical error data can be altered or lost during system restarts, and limited resources for Post Production Repair (PPR) are not optimized for effective error handling.
Innovation Solution
Implementing instructions that store comprehensive error data in non-volatile memory, allowing it to persist across system restarts, and configuring error handling policies to prioritize repairs based on historical data and risk predictions, optimizing the use of PPR resources and reducing downtime.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If error data is stored in volatile memory (MCA registers) during normal operation, then the system can quickly access and process error information, but the data is lost during system restarts
Solution Approach 1:
The patent applies preliminary action by capturing and storing error data in non-volatile memory (NVM) immediately when an error is detected, before the system restart occurs. This ensures the error information is preserved across power cycles and restart events, eliminating the need to rely on volatile MCA registers that would be lost during restart.
Solution Approach 2:
The patent introduces non-volatile memory as an intermediary between the volatile MCA registers and the ultimate error logging destination. The NVM acts as a buffer that temporarily holds error data during system restarts, allowing the error information to survive the transition from volatile to persistent storage without loss.
2Reliability
If the system enters System Management Mode (SMM) to handle errors, then error data can be retrieved from MCA registers, but the operating system execution is suspended
Solution Approach 1:
The patent extracts the error data capture function from the SMM error handling path and implements it directly in the error detection logic. By capturing error data in NVM at the point of detection, the system eliminates the need to suspend OS execution and transfer to SMM for basic error data retrieval, allowing SMM to focus only on critical uncorrectable error responses.
Solution Approach 2:
The patent implements self-service by enabling the error detection mechanism to autonomously capture and store error data in non-volatile memory without requiring intervention from the SMM or OS. The error handling logic independently manages the critical function of preserving error information, reducing dependency on system mode transitions.
3Loss of information
If BMC retrieves error data from MCA registers, then error logging is enabled, but the data may be altered or lost during restarts
Solution Approach 1:
The patent applies preliminary action by having the error detection logic capture and store error data in non-volatile memory before the BMC needs to retrieve it. This preliminary capture ensures the error data is already preserved in a restart-resistant location, eliminating the race condition where BMC retrieval might occur after a restart has corrupted the MCA register data.
Solution Approach 2:
The patent introduces non-volatile memory as an intermediary layer between the MCA registers and the BMC retrieval process. The NVM preserves error data through restart events, allowing the BMC to reliably retrieve accurate error information without the data being altered or lost during the restart transition.
4Ease of repair
If Post Production Repair (PPR) resources are limited, then repair costs are controlled, but effective error handling and repair prioritization become difficult
Solution Approach 1:
The patent implements feedback by capturing comprehensive error data including error type, memory address, and contextual information, then using this feedback to inform repair prioritization decisions. The stored error information provides the basis for analyzing error patterns and making data-driven decisions about which repairs to perform first with limited PPR resources.
Solution Approach 2:
The patent applies preliminary action by capturing and analyzing error data in advance, before repair resources need to be allocated. By having comprehensive error information already stored and analyzed, the system can pre-determine repair priorities and prepare repair plans, eliminating time-consuming analysis during the repair decision-making process.
Data Source
AI summary
A system, method and apparatus to optimize repair in a memory module based on hardware errors identified by microprocessors and a configurable error handling policy. For example, the error handling policy can have a configuration file identifying an amount of repair resources available in the memory module as manufactured. Repair status data can be stored in the memory module to determine repair resources currently available for repair. Further, the error handling policy can be configured with a list of high risk memory addresses prioritized for repair. The list can be used to schedule proactive repair in response to memory errors that would otherwise not be repaired during a typical restarting of the computer system having the memory module.


