Memory Error Processing for Non-Mirror Scrub Success Errors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current memory error processing in server systems does not effectively manage corrected errors, leading to unreliable memory health and performance, as immediate offline actions can cause fragmented memory and trigger hardware RAS characteristics, affecting system reliability and availability.
Innovation Solution
A memory error processing method that differentiates between non-mirror scrub success errors and other errors, taking memory pages offline only when non-mirror scrub success errors reach a threshold, thereby reducing system performance impact and improving compatibility between hardware and software RAS technologies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If immediate offline action is taken when corrected errors occur, then memory health is protected, but system performance is affected and fragmented memory is generated
Solution Approach 1:
The patent applies dynamics by transitioning from immediate offline action to threshold-based offline action. The system dynamically adjusts its response based on the accumulation of corrected errors over time, only taking memory pages offline when the error count reaches a predefined threshold M, thereby balancing memory health protection with system performance maintenance.
Solution Approach 2:
The patent changes the parameter from immediate response to threshold-based response. By introducing a threshold M (where M > 1) for the quantity of corrected errors, the system modifies its behavior from immediate offline action to conditional offline action, reducing unnecessary performance impact while maintaining reliability.
2Reliability
If immediate offline action is taken when corrected errors occur, then memory health is protected, but hardware RAS characteristics are triggered
Solution Approach 1:
The system dynamically adjusts its offline action trigger point from immediate to threshold-based. By requiring M corrected errors to accumulate before triggering offline action, the system avoids prematurely activating hardware RAS characteristics while still maintaining memory health protection.
Solution Approach 2:
The patent applies beforehand cushioning by introducing a buffer (threshold M) between error occurrence and offline action. This cushioning prevents immediate triggering of hardware RAS characteristics, allowing the system to absorb a certain number of corrected errors before taking action that would activate hardware RAS mechanisms.
3Productivity
If corrected errors are not processed, then system performance is maintained, but memory health and RAS reliability are affected
Solution Approach 1:
The patent implements feedback by continuously monitoring and counting corrected errors in memory pages. The system provides feedback through the error counter, which tracks the accumulation of corrected errors over time, enabling informed decisions about when to take offline action to balance performance and reliability.
Solution Approach 2:
The system performs preliminary action by counting and monitoring corrected errors before taking offline action. The error counter accumulates data in advance, allowing the system to prepare for and timing the offline action optimally when the threshold is reached, rather than reacting immediately or never acting.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
This application discloses a memory error processing method and apparatus, relates to the field of computer technologies, and helps improve RAS of memory. The method is applied to a computer apparatus, and the method may include: obtaining first error description information, where the first error description information is used to describe a type of an error that occurs in a first memory page; determining, based on the first error description information, that the error that occurs in the first memory page is a non-mirror scrub success error of corrected errors; and in response to the determining, taking the first memory page offline when a quantity of times that the non-mirror scrub success error occurs in the first memory page reaches M, where M is an integer greater than 1.