ML Processor Bank Swapping for Persistent Fault Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The detection of persistent faults in processors and memories is difficult and time-consuming, especially in special-purpose AI and ML systems where faults can lead to incorrect classifications or detections, and existing methods require high software overhead or complete hardware replication.

Innovation Solution

The system employs bank swapping operations by rotating or swapping existing hardware between different runs to detect persistent faults, using minimal additional hardware and without significant software changes, allowing for the detection of both persistent and transient faults.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If bank swapping operations are implemented to detect persistent faults, then fault detection capability is improved, but device complexity increases

Engineering Contradiction:
Improvefault detection capabilityVSAvoidhardware complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies universality by making existing hardware components serve dual purposes: normal processing operations and fault detection operations. The same processing units and memory banks used for computing are swapped and reused for detection purposes, eliminating the need for separate dedicated detection hardware and thereby reducing overall device complexity while maintaining fault detection capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses copying by creating virtual copies of processing operations through bank swapping. Instead of physically replicating entire hardware systems for detection, the method creates computational copies by swapping memory banks and processing units, allowing fault detection through comparison of results from these virtual copies with minimal additional hardware

Inventive Principle:
Principle #26Copying

2Measurement precision

If traditional fault detection methods are used, then measurement precision is improved, but loss of time increases

Engineering Contradiction:
Improvefault detection accuracyVSAvoiddetection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements continuity of useful action by performing fault detection operations during normal processor idle cycles and between computational tasks. The bank swapping operations are executed during periods when processing units would otherwise be unavailable, ensuring that fault detection occurs continuously without interrupting productive computational work, thereby maintaining detection accuracy while minimizing time loss

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent applies preliminary action by performing fault detection operations in advance during idle cycles before critical computations are executed. By swapping memory banks and verifying processor health proactively during available time windows, the system ensures accurate fault detection is completed before it could impact time-sensitive operations

Inventive Principle:
Principle #10Preliminary action

3Reliability

If complete hardware replication is used for fault detection, then reliability is improved, but device complexity and cost increase

Engineering Contradiction:
Improvefault detection capabilityVSAvoidhardware replication
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent eliminates the need for complete hardware replication by making existing hardware universal. The same processing units, memory banks, and interconnect structures are used for both normal computation and fault detection through time-multiplexed bank swapping operations, achieving reliable fault detection without duplicating hardware

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies dynamics by making the hardware configuration flexible and reconfigurable through bank swapping. Instead of static replicated hardware, the system dynamically reassigns memory banks and processing units between computation and detection modes, allowing a single hardware instance to reliably perform multiple functions through temporal reconfiguration

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12265444B2Systems and methods for detection of persistent faults in processing units and memory
Publication Date: 2025.04.01 NXP BV
  • US12265444B2 patent drawing
  • US12265444B2 patent drawing
  • US12265444B2 patent drawing

AI summary

Systems and methods for detection of persistent faults in processing units and memory have been described. In an illustrative, non-limiting embodiment, a Machine Learning (ML) processor includes one or more registers, and a data moving circuit coupled to the one or more registers. The data moving circuit can be configured to select, based upon a first value stored in the one or more registers, an original one of a plurality of parallel handling circuits within the ML processor to obtain an original data processing result. The data moving circuit can also be configured to select, based upon a second value stored in the one or more registers, an alternative one of the plurality of parallel handling circuits to obtain an alternative data processing result that, upon comparison with the original data processing result, provides an indication of a persistent fault in the ML processor.