Memory Fault Detection Using Correctable Error Pattern Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing memory fault detection methods in servers rely solely on the comparison of the total number of correctable errors with a threshold, which may not accurately reflect the memory's fault status, leading to unreliable and inaccurate fault analysis.

Innovation Solution

A method that uses the number of times of correctable errors as a triggering condition, calculating target detection parameters based on error information to comprehensively analyze the relationship between correctable errors and faults, including determining hard and soft error types, memory units with errors, and their frequencies, to improve fault detection accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the total number of correctable errors is used as the sole triggering condition for fault warnings, then the detection process is simple and quick, but the fault analysis reliability and accuracy deteriorate

Engineering Contradiction:
Improvedetection speedVSAvoidfault analysis reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the fault detection process into two distinct phases: a rapid monitoring phase that tracks the total number of correctable errors, and a detailed analysis phase that examines error types, frequencies, and patterns. This segmentation allows the system to maintain quick detection response while ensuring reliable fault analysis through comprehensive parameter examination.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by using the total number of correctable errors as a triggering condition that initiates further detailed analysis. This preliminary threshold-based detection quickly identifies potential issues, then pre-configures the system to perform comprehensive error analysis including error types, frequencies, and patterns, ensuring both speed and reliability.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If the total number of correctable errors is used as the sole triggering condition, then the detection method is simple to implement, but the fault detection accuracy deteriorates

Engineering Contradiction:
Improvedetection method complexityVSAvoidfault detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The detection methodology is segmented into a simple triggering mechanism based on total correctable error counts, and a comprehensive analysis component that examines error types, frequencies, and patterns. This segmentation maintains implementation simplicity while achieving high detection accuracy through multi-dimensional error characterization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary analysis layer between the simple error count threshold and the final fault detection result. This intermediary layer processes error types, frequencies, and patterns to bridge the gap between simple triggering and accurate fault detection, improving measurement precision without significantly increasing implementation complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If only the cumulative number of correctable errors is considered, then the monitoring process is efficient and fast, but the ability to reflect actual memory fault situations deteriorates

Engineering Contradiction:
Improvemonitoring efficiencyVSAvoidmemory fault situation information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The monitoring process is segmented into efficient aggregate tracking of correctable error totals, and detailed extraction of error characteristics including types, frequencies, and patterns. This segmentation maintains monitoring efficiency while preventing information loss about actual memory fault situations through comprehensive error characterization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds another dimension to error monitoring by examining not just the cumulative number of correctable errors, but also error types, frequencies, and patterns. This dimensional expansion transforms one-dimensional error counting into multi-dimensional error analysis, preserving rich information about memory fault situations while maintaining monitoring efficiency through structured data collection.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12608259B2Method, and device for detecting memory fault, medium and server
Publication Date: 2026.04.21 INSPUR SUZHOU INTELLIGENT TECH CO LTD
  • US12608259B2 patent drawing
  • US12608259B2 patent drawing
  • US12608259B2 patent drawing

AI summary

This solution monitors the operating status of the memory, acquires the error information of the memory, and determines the number of times of the correctable error that occurs in the memory based on the error information; when the number of times of the correctable error reaches the preset trigger number of times, calculates the target detection parameter based on the error information, and determines the fault detection result of the memory based on the target detection parameter. It may be seen that the present application only uses the number of times of the correctable errors as a triggering condition. After this triggering condition is triggered, the target detection parameters will be further calculated based on the error information to conduct in-depth analysis and understanding of the relationship between the number of the correctable errors that occur in the memory and the faults.