Multi-PMU Error Detection and Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting memory errors in computer systems are inadequate when errors from multiple physical memory units (PMUs) combine, even if no single PMU exceeds its error threshold, potentially leading to uncorrectable errors that require system halts.

Innovation Solution

A system and method that detect errors in data output from multiple PMUs, log occurrences, and flag symbols with the highest error counts, allowing evasive or corrective actions to prevent uncorrectable errors by using spare or redundant PMUs, even if individual error frequencies do not exceed predetermined thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If error checking is performed only on individual PMUs against predetermined thresholds, then the detection method is simple and computationally efficient, but errors from multiple PMUs combining are not detected

Engineering Contradiction:
Improvedetection of combined errors from multiple PMUsVSAvoiderror detection mechanism
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the error detection process into two distinct stages: (1) individual PMU error detection against predetermined thresholds, and (2) combined error detection by analyzing error patterns across multiple PMUs. This segmentation allows the system to maintain simple individual checks while adding a layer of综合分析 for combined errors, resolving the contradiction between detection reliability and mechanism complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary analysis layer that processes error data from multiple PMUs before final error determination. This intermediary stage correlates errors across PMUs and identifies combined error patterns, enabling detection of multi-PMU errors without requiring complete redesign of the fundamental threshold-based detection mechanism.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the error threshold is set low to detect more errors, then more errors are caught, but false positives increase and system performance degrades

Engineering Contradiction:
Improveerror detection coverageVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent employs dynamic parameter adjustment by setting different error thresholds for different PMUs based on their individual error characteristics and usage patterns. This allows the system to optimize detection sensitivity for each PMU, catching more errors where appropriate while maintaining high performance where error rates are naturally low, thus resolving the contradiction between detection coverage and system performance.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If corrective action is taken immediately upon detecting errors from multiple PMUs, then uncorrectable errors are prevented, but system interruptions increase

Engineering Contradiction:
Improveprevention of uncorrectable errorsVSAvoidsystem interruptions
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary monitoring and correlation of errors from multiple PMUs before triggering corrective action. By tracking error patterns over time and identifying trends that suggest impending combined errors, the system can take preventive measures proactively, reducing the frequency of actual system interruptions while still preventing uncorrectable errors from occurring.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7290185B2Methods and apparatus for reducing memory errors
Publication Date: 2007.10.30 MARVELL ASIA PTE LTD
  • US7290185B2 patent drawing
  • US7290185B2 patent drawing
  • US7290185B2 patent drawing

AI summary

In a first aspect, a first method is provided for reducing memory errors. The first method includes the steps of (1) detecting at least one error in data output from a first physical memory unit (PMU) of a memory; (2) detecting at least one error in data output from a second PMU of the memory; and (3) setting a bit indicating respective data output from a plurality of PMUs includes errors. Numerous other aspects are provided.