Network-on-Chip Fault Detection via Autonomous Scrubbing Transactions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In systems-on-chip (SoCs) with network-on-chip (NoC) designs, especially those in critical environments like automotive or aviation, there is a need for advanced fault detection to prevent and report potential failures before they occur, especially during important tasks, as traditional methods are inadequate for anticipating and managing failures in real-time.

Innovation Solution

The implementation of a fault detection system within the NoC that includes master and slave scrubbing blocks injecting scrubbing transactions with special bits to detect and report faults, utilizing a fault controller to generate signals for fail-safe transitions, and incorporating autonomous scrubbing operations to ensure all paths are tested regularly, allowing for proactive fault detection and reporting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional fault detection methods are used in NoC systems, then the system structure remains simple, but the ability to detect and report potential failures before they occur is insufficient

Engineering Contradiction:
Improvefault detection capabilityVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements scrubbing blocks that continuously perform preliminary fault detection actions by injecting test transactions into the NoC fabric before actual failures occur. These scrubbing operations proactively identify potential faults in communication paths, allowing the system to detect and report issues before they impact normal operation, thus resolving the contradiction between early fault detection and system complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary scrubbing blocks and fault detection logic that act as mediators between the NoC components and the fault detection system. These intermediaries inject test transactions, monitor responses, and generate fault indicators without disrupting normal data flow, enabling enhanced fault detection capability while maintaining relatively simple system integration.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If no advanced fault detection is implemented, then the system operates smoothly during critical tasks, but potential failures cannot be anticipated before they occur

Engineering Contradiction:
Improvefail-safe transition capabilityVSAvoiddetection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The scrubbing blocks continuously perform preliminary detection actions by injecting test transactions into all NoC communication paths before critical tasks execute. This continuous preliminary monitoring ensures that potential failures are detected and fault indicators are generated before they can impact system functionality, enabling timely fail-safe transitions without interrupting critical operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous scrubbing operations that relentlessly monitor all NoC communication paths without interruption. This continuous detection action ensures that faults are identified as soon as they occur, maintaining constant awareness of system health and enabling immediate response to potential failures, thus resolving the contradiction between detection timing and system reliability.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If comprehensive fault detection logic is added to detect all potential failures, then detection coverage is improved, but the complexity of detecting and measuring faults increases

Engineering Contradiction:
Improvefault detection precisionVSAvoidfault detection complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the fault detection function into distributed scrubbing blocks placed at strategic locations throughout the NoC fabric. Each scrubbing block independently monitors specific communication paths by injecting and tracking test transactions, dividing the complex task of comprehensive fault detection into manageable segments that can be implemented and analyzed separately, thus improving detection precision while controlling complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses test transactions that copy the structure and format of normal data transactions but contain identifiable markers for tracking. These copied transactions traverse the same NoC paths as actual data, allowing the detection logic to monitor fault conditions without requiring separate dedicated test infrastructure, thereby improving detection coverage while minimizing additional complexity.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11294757B2System and method for advanced detection of failures in a network-on-chip
Publication Date: 2022.04.05 ARTERIS INC
  • US11294757B2 patent drawing
  • US11294757B2 patent drawing

AI summary

System and method are disclosed to detect potential failures in a network-on-chip (NoC) before the potential failures happen. The system tests connectivity from a master to all slaves by sending scrub transactions to test all paths. The scrub transactions are identified using a scrub bit. The scrub transactions are generated at a master scrubbing block/unit and terminated at a slave scrubbing block/unit. The slave scrubbing block sends scrub responses to the scrub transactions along the response path. The scrub responses to the scrub transactions are generated at the slave scrubbing block and terminated at the master scrubbing block. This allows detection of potential failures, which are reported to a system monitor. If a potential failure is detected, the system transitions to a fail-safe mode before the failure occurs.