Microprocessor Fault Detection via On-Line Testing and Checkpointing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As silicon technologies advance into the nanometer regime, concerns about transistor reliability increase due to device scaling, leading to issues like gate oxide wear-out, transistor infant mortality, and manufacturing defects that escape testing, which can compromise component yield and lifetime.

Innovation Solution

A mechanism is provided to protect microprocessor pipelines and on-chip memory systems from silicon defects using area-frugal on-line testing techniques combined with system-level checkpointing, enabling speculative computational epochs for verifying hardware integrity and rolling back to a known-good state in case of defects, allowing continued operation in a degraded performance mode.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If device scaling is pursued to improve transistor density and performance, then productivity increases, but reliability deteriorates due to gate oxide wear-out, transistor infant mortality, and manufacturing defects

Engineering Contradiction:
Improvetransistor densityVSAvoidtransistor reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements preliminary action by performing on-line testing and microarchitectural checkpointing before defects can cause system failure. The system continuously monitors hardware integrity during operation and rolls back to checkpointed states if defects are detected, preventing reliability issues from compromising the scaled transistors' productivity benefits

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary testing mechanism that mediates between the scaled transistors and system operation. The on-line testing infrastructure acts as a mediator to detect defects early, while checkpointing provides intermediate recovery points, allowing the system to maintain productivity from scaled devices without sacrificing reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If aggressive burn-in testing is applied to eliminate infant mortality failures, then reliability improves, but device yield deteriorates due to thermal run-away effects destroying robust devices

Engineering Contradiction:
Improvetransistor reliabilityVSAvoiddevice yield
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces expensive, destructive burn-in testing with a cheaper, non-destructive on-line testing approach. Instead of subjecting devices to extreme stress that can destroy robust transistors, the system uses continuous monitoring during normal operation to detect defects, maintaining high yield while achieving reliability goals

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The system performs self-service reliability validation through on-line testing and self-diagnosis. The microprocessor automatically monitors its own hardware integrity during operation, eliminating the need for external aggressive burn-in testing that reduces yield, while still identifying and isolating defective components

Inventive Principle:
Principle #25Self-service

3Reliability

If manufacturing testing time is increased to detect defects, then reliability improves, but productivity deteriorates due to extended test duration affecting manufacturing throughput

Engineering Contradiction:
Improvedefect detectionVSAvoidmanufacturing throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements periodic on-line testing during microprocessor operation rather than requiring extended continuous testing during manufacturing. The system performs testing at periodic intervals during normal operation, allowing manufacturing throughput to maintain high productivity while still achieving comprehensive defect detection for improved reliability

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system maintains continuity of useful action by performing testing during normal microprocessor operation rather than requiring separate manufacturing test phases. The on-line testing occurs continuously or periodically while the processor is operational, eliminating the trade-off between testing duration and manufacturing throughput

Inventive Principle:
Principle #20Continuity of useful action

4Reliability

If on-line testing and checkpointing are implemented to detect silicon defects, then reliability improves, but area cost increases due to additional testing infrastructure

Engineering Contradiction:
Improvesystem reliabilityVSAvoidchip area
Core Design Contradiction:
ReliabilityVSArea of stationary object

Solution Approach 1:

The patent applies partial action by implementing on-line testing and checkpointing selectively in critical microarchitectural components rather than throughout the entire chip. This targeted approach achieves sufficient reliability improvement while minimizing the area overhead associated with testing infrastructure

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system segments the microprocessor into testable modules with individual checkpointing capability. By dividing the architecture into separable units that can be independently tested and rolled back, the patent reduces the overall area cost compared to a monolithic testing approach, as each segment requires only local testing resources

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8341473B2Microprocessor and method for detecting faults therein
Publication Date: 2012.12.25 THE RGT UNIV OF MICHIGAN
  • US8341473B2 patent drawing
  • US8341473B2 patent drawing
  • US8341473B2 patent drawing

AI summary

A microprocessor has a silicon area comprising a plurality of transistors implemented on the silicon area and a fault detection circuit occupying less than 20% of the silicon area and configured to detect faults at runtime in at least 80% of the plurality of transistors.