Fault-Aware Memory Error Mitigation via Hardware Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current memory systems face significant challenges in efficiently identifying and mitigating uncorrectable errors (UEs) in high bandwidth memory (HBM) embedded in processors, leading to system failures and high costs, as existing error correction methods are often unreliable and inefficient due to the inability to accurately determine the underlying fault causing the error.

Innovation Solution

A fault-aware analysis system that correlates hardware configuration data with historical error patterns to predict the specific hardware element causing a detected UE, enabling targeted corrective actions such as post-package repair, adaptive double device data correction, or page offlining, thereby avoiding future UEs and providing a reliable memory health assessment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional error correction methods are used, then correctable errors can be corrected, but uncorrectable errors cannot be reliably identified or mitigated

Engineering Contradiction:
Improveerror correction reliabilityVSAvoidfault identification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary fault-aware analysis by correlating hardware configuration data with historical error patterns before attempting error correction. This preliminary identification of the specific hardware element causing the error enables more reliable and targeted corrective actions, rather than attempting generic error correction methods that may fail for uncorrectable errors.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms by continuously monitoring error patterns and using this information to improve future error mitigation decisions. Historical error data is fed back into the analysis system to refine predictions about which hardware elements are causing errors, thereby improving both reliability of correction and precision of fault identification over time.

Inventive Principle:
Principle #23Feedback

2Productivity

If generic repair actions are taken without identifying the specific fault, then repair may be attempted quickly, but the repair is often unreliable or inefficient

Engineering Contradiction:
Improverepair speedVSAvoidrepair success rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

Before executing repair actions, the system performs preliminary analysis to identify the specific hardware element causing the error. This ensures that repair actions are targeted and appropriate for the actual fault condition, improving both the speed and reliability of repairs by avoiding unnecessary or incorrect repair attempts.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies local quality by tailoring repair actions to the specific hardware element identified as the error source. Rather than applying generic repair methods uniformly, the system selects and executes repair actions that are specifically suited to the identified fault, thereby improving repair success rates while maintaining efficient repair processes.

Inventive Principle:
Principle #3Local quality

3Stability of the object's composition

If the entire processor system is taken down due to a memory error, then system stability is maintained, but system availability and productivity are severely impacted

Engineering Contradiction:
Improvesystem stabilityVSAvoidsystem availability
Core Design Contradiction:
Stability of the object's compositionVSProductivity

Solution Approach 1:

The system extracts or isolates the specific faulty hardware element from the overall memory system through precise identification and targeted repair actions. By addressing only the specific failing component rather than taking down the entire system, the system maintains stability of the overall architecture while preserving availability and productivity of the processor system.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If comprehensive fault analysis is performed to identify the specific error cause, then repair accuracy is improved, but analysis time and system complexity increase

Engineering Contradiction:
Improvefault identification accuracyVSAvoidanalysis system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system introduces an intermediary fault-aware analysis layer that correlates hardware configuration data with historical error patterns. This intermediary analysis mechanism bridges the gap between raw error detection and precise fault identification, improving accuracy without requiring direct complex analysis of the entire memory system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary correlation analysis between hardware configurations and error patterns before detailed fault identification is needed. This preliminary action prepares the system with pre-processed information that speeds up subsequent fault identification, reducing the overall complexity and time required for comprehensive analysis.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240241778A1In-system mitigation of uncorrectable errors based on confidence factors, based on fault-aware analysis
Publication Date: 2024.07.18 INTEL CORP
  • US20240241778A1 patent drawing
  • US20240241778A1 patent drawing
  • US20240241778A1 patent drawing

AI summary

A system (204) can respond to detection of an uncorrectable error (UE) (254) in memory (246) based on fault-aware analysis. The fault-aware analysis enables the system (204) to generate a determination of a specific hardware element of the memory (246) that caused the detected UE (254). In response to detection of a UE (254), the system (204) can correlate a hardware configuration (256) of the memory (246) device with historical data indicating memory (246) faults for hardware elements of the hardware configuration (256). Based on a determination of the specific component that likely caused the UE (254), the system (204) can issue a corrective action for the specific hardware element based on the determination.