Memory Error Location Determination via Address Decomposition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods fail to accurately determine the specific location of memory errors, especially when only one memory correctable error occurs or when memory Patrol Scrub downgrades to a correctable error, making it difficult to identify the affected memory stick.
Innovation Solution
A method involving acquiring a memory error correction log file, extracting relevant information, and performing calculations to determine the CPU, memory controller, channel, and memory stick locations using specific scripts and registers to pinpoint the error location.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the BIOS sends a log to the BMC sel log only when the memory correctable error reaches the preset threshold, then the log recording threshold is met, but the specific memory stick location cannot be determined when only one error occurs
Solution Approach 1:
The patent introduces an intermediary mechanism (the detailed log recording method using memory system address decomposition) between the error detection and location identification. By breaking down the memory system address into CPU location, memory controller location, channel location, and memory stick location components, the system can precisely identify the affected memory stick even when only one error occurs, without requiring threshold-based log recording
Solution Approach 2:
The patent segments the memory system address into multiple hierarchical components: CPU location, memory controller location, channel location, and memory stick location. This segmentation allows precise identification of the error location at the memory stick level, resolving the contradiction between reliable error detection and precise location measurement
2Quantity of substance
If the BIOS does not record memory Patrol Scrub UCE Downgrades to CE errors, then the logging threshold is maintained, but the specific memory stick location cannot be determined
Solution Approach 1:
The patent introduces an intermediary mechanism that processes memory system addresses to extract precise location information. By decomposing the memory system address into hierarchical components, the system can identify memory stick locations for all error types including Patrol Scrub downgrades, without requiring increased log quantity or threshold adjustments
Solution Approach 2:
The patent segments the memory error logging into hierarchical location components. This segmentation enables precise tracking of memory stick locations for various error types, transforming the inability to locate Patrol Scrub downgrades into a solvable problem through address decomposition
3Quantity of substance
If multiple memory sticks are connected to a channel, then the memory system capacity is increased, but it becomes difficult to identify which specific memory stick has the error
Solution Approach 1:
The patent segments the memory system address into hierarchical components including CPU location, memory controller location, channel location, and memory stick location. This segmentation enables precise identification of the affected memory stick even when multiple sticks are connected to the same channel, as each memory stick has a unique location identifier in the decomposed address structure
Solution Approach 2:
The patent adds dimensional breakdown to the memory address identification process. By decomposing the address into multiple dimensional components (CPU, controller, channel, stick), the system can precisely locate errors in multi-stick configurations, transforming a one-dimensional address into a multi-dimensional location map
Data Source
AI summary
A method for determining a location where a memory error occurs comprises acquiring a memory error correction log file which records the error, and extracting a memory address, a MISC register value and an error type corresponding to the error from the log file; when the amount of memory sticks is more than 1, calculating and obtaining a memory system address corresponding to the error according to the memory address, the MISC register value and the error type; calculating the CPU location corresponding to the error and the memory controller location in a local proxy according to the memory system address; calculating a channel location and a channel address corresponding to the error according to the memory system address, the CPU location and the memory controller location; and calculating a memory stick location corresponding to the error according to the channel location and the channel address.


