Memory Error Processing for Non-Mirror Scrub Success Errors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current memory error processing in server systems does not effectively manage corrected errors, leading to unreliable memory health and performance, as immediate offline actions can cause fragmented memory and trigger hardware RAS characteristics, affecting system reliability and availability.

Innovation Solution

A memory error processing method that differentiates between non-mirror scrub success errors and other errors, taking memory pages offline only when non-mirror scrub success errors reach a threshold, thereby reducing system performance impact and improving compatibility between hardware and software RAS technologies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If immediate offline action is taken when corrected errors occur, then memory health is protected, but system performance is affected and fragmented memory is generated

Engineering Contradiction:
Improvememory healthVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies dynamics by transitioning from immediate offline action to threshold-based offline action. The system dynamically adjusts its response based on the accumulation of corrected errors over time, only taking memory pages offline when the error count reaches a predefined threshold M, thereby balancing memory health protection with system performance maintenance.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter from immediate response to threshold-based response. By introducing a threshold M (where M > 1) for the quantity of corrected errors, the system modifies its behavior from immediate offline action to conditional offline action, reducing unnecessary performance impact while maintaining reliability.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If immediate offline action is taken when corrected errors occur, then memory health is protected, but hardware RAS characteristics are triggered

Engineering Contradiction:
Improvememory healthVSAvoidhardware RAS characteristics
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The system dynamically adjusts its offline action trigger point from immediate to threshold-based. By requiring M corrected errors to accumulate before triggering offline action, the system avoids prematurely activating hardware RAS characteristics while still maintaining memory health protection.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies beforehand cushioning by introducing a buffer (threshold M) between error occurrence and offline action. This cushioning prevents immediate triggering of hardware RAS characteristics, allowing the system to absorb a certain number of corrected errors before taking action that would activate hardware RAS mechanisms.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Productivity

If corrected errors are not processed, then system performance is maintained, but memory health and RAS reliability are affected

Engineering Contradiction:
Improvesystem performanceVSAvoidmemory health
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback by continuously monitoring and counting corrected errors in memory pages. The system provides feedback through the error counter, which tracks the accumulation of corrected errors over time, enabling informed decisions about when to take offline action to balance performance and reliability.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary action by counting and monitoring corrected errors before taking offline action. The error counter accumulates data in advance, allowing the system to prepare for and timing the offline action optimally when the threshold is reached, rather than reacting immediately or never acting.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3916557B1Memory error processing method and device
Publication Date: 2024.02.14 XFUSION DIGITAL TECH CO LTD
  • EP3916557B1 patent drawingFigure 1~2
  • EP3916557B1 patent drawingFigure 3
  • EP3916557B1 patent drawingFigure 4

AI summary

This application discloses a memory error processing method and apparatus, relates to the field of computer technologies, and helps improve RAS of memory. The method is applied to a computer apparatus, and the method may include: obtaining first error description information, where the first error description information is used to describe a type of an error that occurs in a first memory page; determining, based on the first error description information, that the error that occurs in the first memory page is a non-mirror scrub success error of corrected errors; and in response to the determining, taking the first memory page offline when a quantity of times that the non-mirror scrub success error occurs in the first memory page reaches M, where M is an integer greater than 1.