Resilient NoC Fault Isolation via Duplicated Logic Units

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current approaches for fault detection and isolation in resilient cache coherent systems are inefficient, particularly in preventing fault propagation and addressing transient and permanent faults in network-on-chip (NoC) systems, and do not effectively handle errors from unit duplication or Error Correcting Codes.

Innovation Solution

The system employs duplicated coherent interconnect units, with a functional logic unit and a corresponding checker logic unit, using clock trees and configurable delays to detect and isolate faults, and an isolation unit to prevent fault propagation by isolating faulty units and resetting the system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If timeout errors at targets or slaves are used to allow recovery after isolation of a network interface unit, then fault isolation is enabled, but a timeout at a master or initiator does not allow recovery and requires a system reset

Engineering Contradiction:
Improvefault isolation capabilityVSAvoidrecovery capability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system segments the fault handling capability by providing different timeout mechanisms for different components: targets/slaves have timeout capability for automatic recovery, while masters/initiators use a different mechanism. This segmentation allows each component to be optimized for its specific role in fault detection and recovery.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An intermediary mechanism is introduced between the master/initiator and the fault isolation process. Instead of direct timeout-based isolation at the master level, the system uses an intermediate recovery process that prevents full system reset while still achieving fault isolation, thereby maintaining both reliability and ease of operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If duplication of all logic units is implemented for fault resilience, then system reliability is improved, but device complexity and power consumption increase

Engineering Contradiction:
Improvefault resilienceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Instead of uniformly duplicating all logic units throughout the system, the patent applies duplication selectively to specific components where it provides the most benefit. This local quality approach maintains fault resilience in critical areas while reducing overall system complexity and power consumption compared to full system duplication.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements partial duplication of logic units rather than complete duplication. This partial action provides sufficient fault resilience for the intended application while avoiding the excessive complexity and resource consumption that would result from duplicating every logic unit in the system.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If timeout approach is used for fault isolation, then fault detection is enabled, but power consumption increases due to requirement of power domain boundary definition

Engineering Contradiction:
Improvefault detection capabilityVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system changes the parameters of fault detection by implementing timeout mechanisms at the logic unit level rather than requiring system-level power domain boundaries. This parameter change enables precise fault detection without the power consumption overhead of creating and managing power domain boundaries, as the timeout approach can operate within existing power domains.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11176297B2Detection and isolation of faults to prevent propagation of faults in a resilient system
Publication Date: 2021.11.16 ARTERIS INC
  • US11176297B2 patent drawing
  • US11176297B2 patent drawing
  • US11176297B2 patent drawing

AI summary

A resilient system implementation in a network-on-ship with at least one functional logic unit and at least one duplicated logic unit. A resilient system and method, in accordance with the invention, are disclosed for detecting a fault or an uncorrectable error and isolating the fault. Isolation of the fault prevents further propagation of the fault throughout the system. The resilient system includes isolation logic or an isolation unit that isolates the fault.