Memory Fault Handling via Soft Post Package Repair

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current memory fault handling methods require a cold reset to initiate hard post package repair (hPPR), which disrupts services and fails to recover severe memory faults in real-time, leading to system breakdowns.

Innovation Solution

A method that analyzes historical fault information to determine the severity of memory faults without cold reset, allowing for immediate fault recovery by replacing faulty rows or banks with redundant ones, using statistical features and thresholds to predict and address memory row or bank faults.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If cold reset is performed to start hard post package repair (hPPR), then memory row fault recovery is achieved, but service continuity is disrupted

Engineering Contradiction:
Improvememory fault recoveryVSAvoidservice continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by performing memory self-check and analyzing fault information before cold reset occurs. The system proactively identifies memory row faults through periodic self-checks and accumulates fault information, then performs soft post package repair (sPPR) in advance to replace faulty rows with redundant rows before the system breakdown forces a cold reset. This preliminary fault recovery action prevents the need for service-disrupting cold resets.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism - the fault information analysis module and soft post package repair (sPPR) functionality - that acts as a mediator between memory faults and cold reset. Instead of directly transitioning from memory fault to cold reset, the system uses this intermediary layer to detect faults, analyze fault information, and perform gradual recovery through sPPR, thereby avoiding the harsh intermediary of cold reset and maintaining service continuity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If memory self-check is performed after cold reset, then memory faults are detected, but fault recovery is delayed until the next cold reset

Engineering Contradiction:
Improvefault detection accuracyVSAvoidfault recovery time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements feedback by continuously accumulating fault information from memory self-checks and using this feedback to trigger soft post package repair (sPPR). The system performs periodic memory self-checks, accumulates fault information, and when thresholds are met or severe faults detected, immediately initiates sPPR to replace faulty rows. This closed-loop feedback mechanism ensures timely fault recovery without waiting for cold reset, reducing fault recovery time while maintaining detection accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies continuity of useful action by making memory fault detection and recovery an ongoing process rather than a periodic post-reset activity. The system continuously performs memory self-checks, accumulates fault information, and can immediately initiate soft post package repair (sPPR) when faults are detected. This continuous monitoring and immediate recovery action eliminates the gaps between cold resets, ensuring uninterrupted fault detection and recovery capabilities.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12014791B2Memory fault handling method and apparatus, device, and storage medium
Publication Date: 2024.06.18 HUAWEI TECH CO LTD
  • US12014791B2 patent drawing
  • US12014791B2 patent drawing
  • US12014791B2 patent drawing

AI summary

The present disclosure provides example memory fault handling method, computer device, and computer-readable storage medium. One example method includes starting fault analysis for a memory at a first moment, where the fault analysis includes obtaining a current fault analysis result of the memory by analyzing historical fault information, the historical fault information includes fault information of the memory accumulated in a historical time period, and the historical time period is a time period before the first moment or a time period before the first moment and including the first moment. Fault recovery is started for the memory based on the current fault analysis result of the memory.