Resilient NoC Fault Isolation via Duplicated Logic Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current approaches for fault detection and isolation in resilient cache coherent systems are inefficient, particularly in preventing fault propagation and addressing transient and permanent faults in network-on-chip (NoC) systems, and do not effectively handle errors from unit duplication or Error Correcting Codes.
Innovation Solution
The system employs duplicated coherent interconnect units, with a functional logic unit and a corresponding checker logic unit, using clock trees and configurable delays to detect and isolate faults, and an isolation unit to prevent fault propagation by isolating faulty units and resetting the system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If timeout errors at targets or slaves are used to allow recovery after isolation of a network interface unit, then fault isolation is enabled, but a timeout at a master or initiator does not allow recovery and requires a system reset
Solution Approach 1:
The system segments the fault handling capability by providing different timeout mechanisms for different components: targets/slaves have timeout capability for automatic recovery, while masters/initiators use a different mechanism. This segmentation allows each component to be optimized for its specific role in fault detection and recovery.
Solution Approach 2:
An intermediary mechanism is introduced between the master/initiator and the fault isolation process. Instead of direct timeout-based isolation at the master level, the system uses an intermediate recovery process that prevents full system reset while still achieving fault isolation, thereby maintaining both reliability and ease of operation.
2Reliability
If duplication of all logic units is implemented for fault resilience, then system reliability is improved, but device complexity and power consumption increase
Solution Approach 1:
Instead of uniformly duplicating all logic units throughout the system, the patent applies duplication selectively to specific components where it provides the most benefit. This local quality approach maintains fault resilience in critical areas while reducing overall system complexity and power consumption compared to full system duplication.
Solution Approach 2:
The system implements partial duplication of logic units rather than complete duplication. This partial action provides sufficient fault resilience for the intended application while avoiding the excessive complexity and resource consumption that would result from duplicating every logic unit in the system.
3Measurement precision
If timeout approach is used for fault isolation, then fault detection is enabled, but power consumption increases due to requirement of power domain boundary definition
Solution Approach 1:
The system changes the parameters of fault detection by implementing timeout mechanisms at the logic unit level rather than requiring system-level power domain boundaries. This parameter change enables precise fault detection without the power consumption overhead of creating and managing power domain boundaries, as the timeout approach can operate within existing power domains.
Data Source
AI summary
A resilient system implementation in a network-on-ship with at least one functional logic unit and at least one duplicated logic unit. A resilient system and method, in accordance with the invention, are disclosed for detecting a fault or an uncorrectable error and isolating the fault. Isolation of the fault prevents further propagation of the fault throughout the system. The resilient system includes isolation logic or an isolation unit that isolates the fault.


