ML Processor Bank Swapping for Persistent Fault Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The detection of persistent faults in processors and memories is difficult and time-consuming, especially in special-purpose AI and ML systems where faults can lead to incorrect classifications or detections, and existing methods require high software overhead or complete hardware replication.
Innovation Solution
The system employs bank swapping operations by rotating or swapping existing hardware between different runs to detect persistent faults, using minimal additional hardware and without significant software changes, allowing for the detection of both persistent and transient faults.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If bank swapping operations are implemented to detect persistent faults, then fault detection capability is improved, but device complexity increases
Solution Approach 1:
The patent applies universality by making existing hardware components serve dual purposes: normal processing operations and fault detection operations. The same processing units and memory banks used for computing are swapped and reused for detection purposes, eliminating the need for separate dedicated detection hardware and thereby reducing overall device complexity while maintaining fault detection capability
Solution Approach 2:
The patent uses copying by creating virtual copies of processing operations through bank swapping. Instead of physically replicating entire hardware systems for detection, the method creates computational copies by swapping memory banks and processing units, allowing fault detection through comparison of results from these virtual copies with minimal additional hardware
2Measurement precision
If traditional fault detection methods are used, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent implements continuity of useful action by performing fault detection operations during normal processor idle cycles and between computational tasks. The bank swapping operations are executed during periods when processing units would otherwise be unavailable, ensuring that fault detection occurs continuously without interrupting productive computational work, thereby maintaining detection accuracy while minimizing time loss
Solution Approach 2:
The patent applies preliminary action by performing fault detection operations in advance during idle cycles before critical computations are executed. By swapping memory banks and verifying processor health proactively during available time windows, the system ensures accurate fault detection is completed before it could impact time-sensitive operations
3Reliability
If complete hardware replication is used for fault detection, then reliability is improved, but device complexity and cost increase
Solution Approach 1:
The patent eliminates the need for complete hardware replication by making existing hardware universal. The same processing units, memory banks, and interconnect structures are used for both normal computation and fault detection through time-multiplexed bank swapping operations, achieving reliable fault detection without duplicating hardware
Solution Approach 2:
The patent applies dynamics by making the hardware configuration flexible and reconfigurable through bank swapping. Instead of static replicated hardware, the system dynamically reassigns memory banks and processing units between computation and detection modes, allowing a single hardware instance to reliably perform multiple functions through temporal reconfiguration
Data Source
AI summary
Systems and methods for detection of persistent faults in processing units and memory have been described. In an illustrative, non-limiting embodiment, a Machine Learning (ML) processor includes one or more registers, and a data moving circuit coupled to the one or more registers. The data moving circuit can be configured to select, based upon a first value stored in the one or more registers, an original one of a plurality of parallel handling circuits within the ML processor to obtain an original data processing result. The data moving circuit can also be configured to select, based upon a second value stored in the one or more registers, an alternative one of the plurality of parallel handling circuits to obtain an alternative data processing result that, upon comparison with the original data processing result, provides an indication of a persistent fault in the ML processor.


