Memory Error Data Preservation via Non-Volatile Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in accurately logging and handling uncorrectable hardware errors in memory systems accessed by microprocessors, as critical error data can be altered or lost during system restarts, and limited resources for Post Production Repair (PPR) are not optimized for effective error handling.

Innovation Solution

Implementing instructions that store comprehensive error data in non-volatile memory, allowing it to persist across system restarts, and configuring error handling policies to prioritize repairs based on historical data and risk predictions, optimizing the use of PPR resources and reducing downtime.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If error data is stored in volatile memory (MCA registers) during normal operation, then the system can quickly access and process error information, but the data is lost during system restarts

Engineering Contradiction:
Improveerror data preservationVSAvoidsystem restart downtime
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent applies preliminary action by capturing and storing error data in non-volatile memory (NVM) immediately when an error is detected, before the system restart occurs. This ensures the error information is preserved across power cycles and restart events, eliminating the need to rely on volatile MCA registers that would be lost during restart.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces non-volatile memory as an intermediary between the volatile MCA registers and the ultimate error logging destination. The NVM acts as a buffer that temporarily holds error data during system restarts, allowing the error information to survive the transition from volatile to persistent storage without loss.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system enters System Management Mode (SMM) to handle errors, then error data can be retrieved from MCA registers, but the operating system execution is suspended

Engineering Contradiction:
Improveerror handling capabilityVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts the error data capture function from the SMM error handling path and implements it directly in the error detection logic. By capturing error data in NVM at the point of detection, the system eliminates the need to suspend OS execution and transfer to SMM for basic error data retrieval, allowing SMM to focus only on critical uncorrectable error responses.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements self-service by enabling the error detection mechanism to autonomously capture and store error data in non-volatile memory without requiring intervention from the SMM or OS. The error handling logic independently manages the critical function of preserving error information, reducing dependency on system mode transitions.

Inventive Principle:
Principle #25Self-service

3Loss of information

If BMC retrieves error data from MCA registers, then error logging is enabled, but the data may be altered or lost during restarts

Engineering Contradiction:
Improveerror data integrityVSAvoiderror handling architecture
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by having the error detection logic capture and store error data in non-volatile memory before the BMC needs to retrieve it. This preliminary capture ensures the error data is already preserved in a restart-resistant location, eliminating the race condition where BMC retrieval might occur after a restart has corrupted the MCA register data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces non-volatile memory as an intermediary layer between the MCA registers and the BMC retrieval process. The NVM preserves error data through restart events, allowing the BMC to reliably retrieve accurate error information without the data being altered or lost during the restart transition.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of repair

If Post Production Repair (PPR) resources are limited, then repair costs are controlled, but effective error handling and repair prioritization become difficult

Engineering Contradiction:
Improverepair resource optimizationVSAvoidrepair decision making time
Core Design Contradiction:
Ease of repairVSLoss of time

Solution Approach 1:

The patent implements feedback by capturing comprehensive error data including error type, memory address, and contextual information, then using this feedback to inform repair prioritization decisions. The stored error information provides the basis for analyzing error patterns and making data-driven decisions about which repairs to perform first with limited PPR resources.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary action by capturing and analyzing error data in advance, before repair resources need to be allocated. By having comprehensive error information already stored and analyzed, the system can pre-determine repair priorities and prepare repair plans, eliminating time-consuming analysis during the repair decision-making process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11726873B2Handling memory errors identified by microprocessors
Publication Date: 2023.08.15 MICRON TECHNOLOGY INC
  • US11726873B2 patent drawing
  • US11726873B2 patent drawing
  • US11726873B2 patent drawing

AI summary

A system, method and apparatus to optimize repair in a memory module based on hardware errors identified by microprocessors and a configurable error handling policy. For example, the error handling policy can have a configuration file identifying an amount of repair resources available in the memory module as manufactured. Repair status data can be stored in the memory module to determine repair resources currently available for repair. Further, the error handling policy can be configured with a list of high risk memory addresses prioritized for repair. The list can be used to schedule proactive repair in response to memory errors that would otherwise not be repaired during a typical restarting of the computer system having the memory module.