Microprocessor Error Data Preservation via SPD Memory
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for recording hardware errors in memory systems accessed by microprocessors face challenges in preserving accurate data, especially for uncorrectable errors, as the data can be altered during system restarts, and existing solutions do not capture comprehensive information necessary for detailed analysis and prediction of future failures.
Innovation Solution
Implementing instructions that store error data in a non-volatile memory region supporting Serial Presence Detect (SPD), allowing for comprehensive data collection and preservation, including error details like memory addresses and configuration parameters, which remain intact even after system restarts, and communicating this data to a Baseboard Management Controller (BMC) for independent storage and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error data is stored in volatile memory (MCA registers) and accessed during system operation, then the data can be retrieved by the operating system or BMC, but the data is lost or altered during system restarts
Solution Approach 1:
The patent applies preliminary action by copying error data from volatile MCA registers to non-volatile memory (such as SPD memory or persistent storage) before the system restart occurs. This ensures that error data is preserved in advance, preventing information loss during the restart process. The BMC or firmware performs this copy operation as part of the error handling routine before system state changes.
2Loss of information
If comprehensive error data including memory addresses and configuration parameters is collected and stored, then detailed error analysis and prediction of future failures is enabled, but the complexity of the error logging system increases
Solution Approach 1:
The patent applies universality by using the existing SPD (Serial Presence Detect) memory, which is already present in memory modules for storing manufacturer data, to also store error information. This multi-functional use of existing hardware resources enables comprehensive error data collection without adding separate dedicated storage components, thereby reducing system complexity while achieving detailed error tracking.
Solution Approach 2:
The BMC (Baseboard Management Controller) serves as an intermediary that collects, processes, and stores comprehensive error data from multiple sources including MCA registers and memory configuration parameters. The BMC consolidates this information into a unified error log, simplifying the complexity by providing a single point of error management rather than requiring direct integration across multiple system components.
3Reliability
If error data is stored in non-volatile memory regions like SPD memory, then data persists across system restarts enabling post-restart analysis, but the available storage capacity in these regions is reduced
Solution Approach 1:
The patent applies local quality by designating specific portions or regions of the SPD memory for error data storage while preserving other regions for their original purposes such as manufacturer specifications and memory configuration. This segmented approach allows error logging functionality to coexist with the original SPD memory functions, maintaining sufficient storage capacity for both purposes without requiring complete dedication of the memory region.
Data Source
AI summary
A system, method and apparatus to record data relevant to hardware errors identified by microprocessors. For example, in response to a hardware error, a microprocessor can store first data about the error in registers in the microprocessor and start to execute instructions configured in firmware and/or in an operating system. Execution of the instructions in response to the hardware error causes the microprocessor to: generating second data about the error based at least in part on the first data in the registers; and store the second data at a location not affected by restarting execution of an operating system in the processor. For example, the execution of the instructions can cause the microprocessor to decode the first data to obtain a temperature of the computing device as part of the second data.


