Hierarchical Fault Masking in SoC Reset Operations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In a system on a chip (SoC) with increasing levels of system integration, hardware faults in processing systems can disrupt the reliable operation of subsystems and the overall SoC, requiring effective fault detection and recovery mechanisms without affecting other subsystems.
Innovation Solution
A hierarchical fault handling system is implemented, where a central fault collection and control system (FCCS) communicates with local FCCSs in subsystems to detect and manage faults. This system performs asynchronous communication of fault indications, software resets, and hardware resets, with masking of fault indications to prevent disruptions during reset processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fault detection and propagation mechanisms are implemented in a highly integrated SoC system, then fault detection capability is improved, but system complexity and the risk of additional fault indications during reset operations increase
Solution Approach 1:
The system is divided into multiple hierarchical levels with central FCCS and local FCCS units. Each level handles fault detection and management independently within its scope, reducing the complexity burden on any single component while maintaining comprehensive fault detection across the entire SoC system.
Solution Approach 2:
The fault masking mechanism acts as an intermediary between fault detection and fault propagation. During reset operations, the masking mechanism temporarily blocks fault indications from propagating upward, preventing reset-related noise from overwhelming the fault detection system while still allowing genuine faults to be detected and reported after reset completion.
2Reliability
If reset operations are performed to recover from faults, then system reliability is improved, but additional fault indications may be generated that disrupt the reset process
Solution Approach 1:
Before initiating reset operations, the system activates a masking mechanism that preemptively blocks fault indications that would normally be generated during reset. This preliminary anti-action prevents the harmful effect of additional fault indications from disrupting the reset process, allowing the reset to complete successfully while maintaining fault detection capability for genuine faults.
3Measurement precision
If comprehensive fault monitoring is implemented across all subsystems, then fault detection precision is improved, but the system becomes more susceptible to fault indication storms during error conditions
Solution Approach 1:
Different parts of the system have different fault monitoring characteristics. Local FCCS units provide detailed fault detection within their respective subsystems, while the central FCCS provides aggregate monitoring. The masking mechanism applies selectively during reset operations, allowing precise local fault detection to continue while preventing the amplification of fault indications into storms during vulnerable periods.
Data Source
AI summary
A central system coupled to a subsystem receives a fault indication associated with a fault in one or more circuits of the subsystem from a local fault collection and control (FCCS) of the subsystem when a software recovery of the fault fails. Based on the received fault indication, the local FCCS and a central FCCS of a central system is masked from additional fault indications from the one or more circuits. The central system then signals the reset of the one or more circuits of the subsystem after the masking of the additional fault indications, wherein the one or more circuits is reset based on the signaling and the additional faults are masked from one or more of the local FCCS and central FCCS during the reset.


