Network-on-Chip Fault Detection via Autonomous Scrubbing Transactions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In systems-on-chip (SoCs) with network-on-chip (NoC) designs, especially those in critical environments like automotive or aviation, there is a need for advanced fault detection to prevent and report potential failures before they occur, especially during important tasks, as traditional methods are inadequate for anticipating and managing failures in real-time.
Innovation Solution
The implementation of a fault detection system within the NoC that includes master and slave scrubbing blocks injecting scrubbing transactions with special bits to detect and report faults, utilizing a fault controller to generate signals for fail-safe transitions, and incorporating autonomous scrubbing operations to ensure all paths are tested regularly, allowing for proactive fault detection and reporting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional fault detection methods are used in NoC systems, then the system structure remains simple, but the ability to detect and report potential failures before they occur is insufficient
Solution Approach 1:
The patent implements scrubbing blocks that continuously perform preliminary fault detection actions by injecting test transactions into the NoC fabric before actual failures occur. These scrubbing operations proactively identify potential faults in communication paths, allowing the system to detect and report issues before they impact normal operation, thus resolving the contradiction between early fault detection and system complexity.
Solution Approach 2:
The patent introduces intermediary scrubbing blocks and fault detection logic that act as mediators between the NoC components and the fault detection system. These intermediaries inject test transactions, monitor responses, and generate fault indicators without disrupting normal data flow, enabling enhanced fault detection capability while maintaining relatively simple system integration.
2Reliability
If no advanced fault detection is implemented, then the system operates smoothly during critical tasks, but potential failures cannot be anticipated before they occur
Solution Approach 1:
The scrubbing blocks continuously perform preliminary detection actions by injecting test transactions into all NoC communication paths before critical tasks execute. This continuous preliminary monitoring ensures that potential failures are detected and fault indicators are generated before they can impact system functionality, enabling timely fail-safe transitions without interrupting critical operations.
Solution Approach 2:
The patent implements continuous scrubbing operations that relentlessly monitor all NoC communication paths without interruption. This continuous detection action ensures that faults are identified as soon as they occur, maintaining constant awareness of system health and enabling immediate response to potential failures, thus resolving the contradiction between detection timing and system reliability.
3Measurement precision
If comprehensive fault detection logic is added to detect all potential failures, then detection coverage is improved, but the complexity of detecting and measuring faults increases
Solution Approach 1:
The patent segments the fault detection function into distributed scrubbing blocks placed at strategic locations throughout the NoC fabric. Each scrubbing block independently monitors specific communication paths by injecting and tracking test transactions, dividing the complex task of comprehensive fault detection into manageable segments that can be implemented and analyzed separately, thus improving detection precision while controlling complexity.
Solution Approach 2:
The patent uses test transactions that copy the structure and format of normal data transactions but contain identifiable markers for tracking. These copied transactions traverse the same NoC paths as actual data, allowing the detection logic to monitor fault conditions without requiring separate dedicated test infrastructure, thereby improving detection coverage while minimizing additional complexity.
Data Source
AI summary
System and method are disclosed to detect potential failures in a network-on-chip (NoC) before the potential failures happen. The system tests connectivity from a master to all slaves by sending scrub transactions to test all paths. The scrub transactions are identified using a scrub bit. The scrub transactions are generated at a master scrubbing block/unit and terminated at a slave scrubbing block/unit. The slave scrubbing block sends scrub responses to the scrub transactions along the response path. The scrub responses to the scrub transactions are generated at the slave scrubbing block and terminated at the master scrubbing block. This allows detection of potential failures, which are reported to a system monitor. If a potential failure is detected, the system transitions to a fail-safe mode before the failure occurs.

