Dynamic Soft-Error Rate Discrimination via Parity-Space Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for distinguishing between soft errors and the onset of hardware degradation in computer systems are limited by high false alarm rates, inability to account for dynamic variations in cosmic neutron flux, and altitude-related acceleration, leading to inefficient detection and excessive hardware replacement.

Innovation Solution

A dynamic Soft-Error Rate Discrimination (SERD) technique using in-situ self-sensing and parity-space detection, which averages and normalizes correctable-error events across multiple memory components to generate residual vectors, and applies a Sequential Probability Ratio Test to differentiate between soft errors and hardware degradation, while being insensitive to dynamic cosmic neutron variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a fixed-threshold SERD technique is used to detect soft errors, then the detection is simple to implement, but it cannot account for dynamic variations in cosmic neutron flux leading to high false alarm rates

Engineering Contradiction:
Improvesimplicity of SERD implementationVSAvoidfalse alarm rate
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent implements dynamic threshold adjustment by continuously monitoring the average error rate across multiple memory components and adapting the threshold accordingly. The threshold is no longer fixed but dynamically responds to changing cosmic neutron flux conditions, reducing false alarms while maintaining detection sensitivity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from multiple memory components to adjust the detection threshold. By monitoring the average error rate across all components and using this information to modify the threshold, the system creates a self-regulating mechanism that adapts to dynamic environmental conditions.

Inventive Principle:
Principle #23Feedback

2Reliability

If the SERD threshold is set high to avoid false alarms during cosmic neutron peaks, then false alarms are reduced, but sensitivity to hardware degradation during cosmic troughs is lost

Engineering Contradiction:
Improvefalse alarm reductionVSAvoidsensitivity to hardware degradation
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The threshold dynamically adjusts based on current cosmic neutron flux conditions inferred from the average error rate across memory components. During cosmic peaks, the threshold increases to reduce false alarms, while during troughs, it decreases to maintain sensitivity to actual hardware degradation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The detection threshold parameter is changed dynamically based on system conditions. The threshold is not a fixed value but varies as a function of the average error rate, allowing optimal sensitivity across different operating conditions.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If a constant-threshold leaky bucket technique is used, then the implementation is simple, but it cannot accommodate altitude variations causing up to 70% acceleration in cosmic neutron flux

Engineering Contradiction:
Improveimplementation simplicityVSAvoidaltitude adaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system performs self-characterization by monitoring its own error rate statistics to infer altitude and cosmic neutron flux conditions. Each system automatically adapts to its specific operating conditions without external intervention, making the solution self-service and universally applicable across different altitudes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The detection parameters are changed based on inferred altitude conditions. By monitoring the average error rate and comparing it against expected values, the system adjusts its threshold behavior to accommodate the 70% flux acceleration at high altitudes versus sea level.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If multiple memory components are monitored to improve detection accuracy, then the ability to distinguish soft errors from hardware degradation improves, but the system complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidnumber of monitored components
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Multiple memory components are merged into a single statistical analysis. Instead of analyzing each component independently, the system combines their error rates into an average, reducing the effective number of separate monitoring streams while maintaining the benefits of multi-component observation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The same monitoring and analysis mechanism is applied universally across all memory components. A single algorithm processes data from multiple components, making the system scalable without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7447957B1Dynamic soft-error-rate discrimination via in-situ self-sensing coupled with parity-space detection
Publication Date: 2008.11.04 ORACLE AMERICAN INC
  • US7447957B1 patent drawing
  • US7447957B1 patent drawing
  • US7447957B1 patent drawing

AI summary

A system that facilitates distinguishing between soft errors and the onset of hardware degradation in a computer system. During operation, the system receives notifications of correctable-error events from a plurality of memory components. The system then averages numbers of correctable-error events from the plurality of memory components to generate an average number of correctable-error events across the plurality of memory components. The system subtracts the number of correctable-error events for a given memory component in a given time interval from the average number of correctable-error events to generate a residual number of correctable-error events for the given memory component in the given time interval.