SoC Fault Controller for Interconnect Queue Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional fault recovery systems in system-on-chips (SoCs) rely heavily on faulty processors to complete ongoing transactions, leading to significant recovery times and reduced SoC utilization due to the unreliable nature of these processors, which can cause system hangs and queue the interconnect.

Innovation Solution

A fault recovery system with a fault controller that receives a time-out signal when a processor fails to execute a transaction, generates control signals to disconnect the processor from the interconnect, and executes the transaction to dequeue the interconnect, thereby managing fault recovery independently of the faulty processor.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the faulty processor is relied upon to complete ongoing transactions, then the fault recovery system can manage recovery, but the recovery time becomes significant and SoC utilization degrades

Engineering Contradiction:
Improvefault recovery managementVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts the transaction completion function from the faulty processor by introducing a fault recovery system that takes over execution of ongoing transactions. When a fault is detected, the fault recovery system interceptors the faulty processor's transactions and completes them independently, removing the unreliable processor from the critical path of transaction completion and enabling faster recovery without waiting for the faulty processor to self-complete transactions

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The fault recovery system acts as an intermediary between the faulty processor and the interconnect. It receives transactions from the faulty processor, manages their completion, and handles the recovery process without requiring the faulty processor to directly complete transactions. This intermediary role allows the system to manage recovery while avoiding the time loss associated with relying on the faulty processor's uncontrolled execution

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the faulty processor is allowed to execute transactions, then transaction completion may be achieved, but the processor may become unresponsive and queue the interconnect

Engineering Contradiction:
Improvetransaction executionVSAvoidsystem stability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements preliminary action by detecting faults before they cause system-wide failures. The fault recovery system continuously monitors processors and, upon detecting a fault, immediately takes over transaction execution and isolates the faulty processor. This preliminary intervention prevents the processor from becoming unresponsive and queuing the interconnect, maintaining system stability while ensuring transaction completion through the fault recovery system's controlled execution

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the faulty processor is disconnected immediately upon fault detection, then system stability is maintained, but ongoing transactions cannot be completed

Engineering Contradiction:
Improvesystem stabilityVSAvoidtransaction completion
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The fault recovery system performs preliminary action by preparing to take over transaction execution before the faulty processor is disconnected. Upon fault detection, it immediately intercepts ongoing transactions, maintains their execution state, and completes them without interruption. This preliminary preparation ensures both system stability (through immediate isolation of the faulty processor) and transaction completion (through the fault recovery system's assumption of execution responsibility)

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11609821B2Method and system for managing fault recovery in system-on-chips
Publication Date: 2023.03.21 NXP USA INC
  • US11609821B2 patent drawing
  • US11609821B2 patent drawing
  • US11609821B2 patent drawing

AI summary

A fault recovery system including a fault controller is disclosed. The fault controller is coupled between a processor and an interconnect, and configured to receive a time-out signal that is indicative of a failure of the processor to execute a transaction after a fault is detected in the processor. The failure in the execution of the transaction results in queuing of the interconnect. Based on the time-out signal, the fault controller is further configured to generate and transmit a control signal to the processor to disconnect the processor from the interconnect. Further, the fault controller is configured to execute the transaction, and in turn, dequeue the interconnect. When the transaction is successfully executed, the fault controller is further configured to generate a status signal to reset the processor, thereby managing a fault recovery of the processor.