GPU Fault Injection Architecture for Precise Soft Error Emulation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for assessing the resilience of graphics processing units (GPUs) to soft errors, such as ionizing radiation, lack control over which memory circuits are affected and when, providing an inaccurate assessment of resilience during application execution.
Innovation Solution
A method for precise state error injection in GPUs, allowing developers to emulate soft errors by halting specific subsystems, injecting errors into specified data storage circuits or RAM bits, and restoring the system to functional mode, enabling transparent assessment of error consequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If ionizing radiation is used to assess GPU resilience, then soft errors can be induced in memory circuits, but control over which specific circuits are affected and when is lost
Solution Approach 1:
The patent introduces an intermediary error injection mechanism that sits between the radiation source and the GPU memory circuits. This intermediary system allows controlled induction of soft errors at specific memory locations and times without requiring actual ionizing radiation, thereby achieving both reliability assessment and precise measurement control
Solution Approach 2:
The patent creates a virtual copy of the radiation-induced error effect through software-controlled error injection. Instead of using physical radiation to induce errors, the system replicates the error conditions through logical bit flips in memory circuits, achieving the same assessment goal with precise control over error location and timing
2Reliability
If conventional radiation testing is used, then soft errors can be induced, but accurate assessment of resilience during specific application execution is prevented
Solution Approach 1:
The patent implements preliminary error injection capabilities that allow errors to be injected at predetermined points during application execution. The system can pre-identify critical execution points and inject errors at those specific moments without interrupting the overall application flow, enabling resilience assessment during actual application runtime
Solution Approach 2:
The patent creates a dynamic error injection system that adapts to the application execution state. The error injection mechanism can dynamically select when and where to inject errors based on the current execution context, allowing continuous application execution while assessing resilience at critical points
3Productivity
If multi-threaded programs are executed, then productivity is improved, but identifying the specific thread affected by an error becomes difficult
Solution Approach 1:
The patent implements a feedback mechanism that tracks the state of each thread and memory circuit throughout execution. When an error is injected or detected, the system provides feedback information about which specific thread and memory circuit are affected, enabling precise error source identification even in multi-threaded environments while maintaining high productivity
Data Source
AI summary
Unavoidable physical phenomena, such as an alpha particle strikes, can cause soft errors in integrated circuits. Materials that emit alpha particles are ubiquitous, and higher energy cosmic particles penetrate the atmosphere and also cause soft errors. Some soft errors have no consequence, but others can cause an integrated circuit to malfunction. In some applications (e.g. driverless cars), proper operation of integrated circuits is critical to human life and safety. To minimize or eliminate the likelihood of a soft error becoming a serious malfunction, detailed assessment of individual potential soft errors and subsequent processor behavior is necessary. Embodiments of the present disclosure facilitate emulating a plurality of different, specific soft errors. Resilience may be assessed over the plurality of soft errors and application code may be advantageously engineered to improve resilience. Normal processor execution is halted to inject a given state error through a scan chain, and execution is subsequently resumed.


