Memory Fault Handling via Soft Post Package Repair
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current memory fault handling methods require a cold reset to initiate hard post package repair (hPPR), which disrupts services and fails to recover severe memory faults in real-time, leading to system breakdowns.
Innovation Solution
A method that analyzes historical fault information to determine the severity of memory faults without cold reset, allowing for immediate fault recovery by replacing faulty rows or banks with redundant ones, using statistical features and thresholds to predict and address memory row or bank faults.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If cold reset is performed to start hard post package repair (hPPR), then memory row fault recovery is achieved, but service continuity is disrupted
Solution Approach 1:
The patent applies preliminary action by performing memory self-check and analyzing fault information before cold reset occurs. The system proactively identifies memory row faults through periodic self-checks and accumulates fault information, then performs soft post package repair (sPPR) in advance to replace faulty rows with redundant rows before the system breakdown forces a cold reset. This preliminary fault recovery action prevents the need for service-disrupting cold resets.
Solution Approach 2:
The patent introduces an intermediary mechanism - the fault information analysis module and soft post package repair (sPPR) functionality - that acts as a mediator between memory faults and cold reset. Instead of directly transitioning from memory fault to cold reset, the system uses this intermediary layer to detect faults, analyze fault information, and perform gradual recovery through sPPR, thereby avoiding the harsh intermediary of cold reset and maintaining service continuity.
2Measurement precision
If memory self-check is performed after cold reset, then memory faults are detected, but fault recovery is delayed until the next cold reset
Solution Approach 1:
The patent implements feedback by continuously accumulating fault information from memory self-checks and using this feedback to trigger soft post package repair (sPPR). The system performs periodic memory self-checks, accumulates fault information, and when thresholds are met or severe faults detected, immediately initiates sPPR to replace faulty rows. This closed-loop feedback mechanism ensures timely fault recovery without waiting for cold reset, reducing fault recovery time while maintaining detection accuracy.
Solution Approach 2:
The patent applies continuity of useful action by making memory fault detection and recovery an ongoing process rather than a periodic post-reset activity. The system continuously performs memory self-checks, accumulates fault information, and can immediately initiate soft post package repair (sPPR) when faults are detected. This continuous monitoring and immediate recovery action eliminates the gaps between cold resets, ensuring uninterrupted fault detection and recovery capabilities.
Data Source
AI summary
The present disclosure provides example memory fault handling method, computer device, and computer-readable storage medium. One example method includes starting fault analysis for a memory at a first moment, where the fault analysis includes obtaining a current fault analysis result of the memory by analyzing historical fault information, the historical fault information includes fault information of the memory accumulated in a historical time period, and the historical time period is a time period before the first moment or a time period before the first moment and including the first moment. Fault recovery is started for the memory based on the current fault analysis result of the memory.


