SoC Self-Recovery Engine for Bit Flip Crash Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems fail to detect and resolve bit flip events during mission mode in SoCs, leading to system crashes and the need for costly returns to the OEM.
Innovation Solution
Implement an ISSR engine with decision and recovery logic to monitor and recover from bit corruption events by increasing supply voltage and restoring the previous system state, allowing continued operation if recovery is successful.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If no detection and recovery mechanism is implemented for bit flip events, then the system structure remains simple, but system reliability deteriorates due to crashes during mission mode
Solution Approach 1:
The ISSR engine enables the system to detect and recover from bit flip events autonomously without external intervention. The engine monitors subsystem crashes, determines bit corruption events, and executes recovery steps automatically, allowing the system to serve itself during failure conditions.
Solution Approach 2:
The system performs preliminary actions by monitoring subsystem status continuously and preparing recovery protocols in advance. When a bit flip event is detected, the ISSR engine has already established the framework for recovery, including voltage adjustment and state restoration mechanisms.
2Productivity
If the system halts upon detecting a bit flip event, then recovery operations can be performed, but productivity deteriorates due to system halt and potential RMA returns
Solution Approach 1:
The ISSR engine changes critical system parameters including supply voltage and operational state to enable recovery. By adjusting voltage levels and restoring previous system states, the engine transforms the crash condition into a recoverable state, allowing continued operation.
Solution Approach 2:
The system implements feedback mechanisms where the ISSR engine continuously monitors subsystem performance and adjusts operations based on detected anomalies. This closed-loop approach enables real-time recovery decisions and maintains productivity by preventing permanent system halt.
3Loss of information
If bit flip events are not detected during mission mode, then the system operates continuously, but loss of information increases due to incorrect data and crashes
Solution Approach 1:
The ISSR engine acts as an intermediary between the subsystems and the overall system control. It monitors subsystem operations, detects bit flip events, and coordinates recovery actions, serving as a mediator that protects data integrity without requiring direct intervention in all subsystem operations.
Data Source
AI summary
Systems and methods for in-system, self-recovery (ISSR) in a system-on-a-chip (SoC) are disclosed for detecting and recovering from a bit corruption event that has caused an SoC subsystem to crash. If a bit flip event is detected through observation of a subsystem crash, ISSR steps are taken in an attempt to correct the issue. If the ISSR steps are successful, then the system is not halted and mission mode operations continue. If the ISSR steps are unsuccessful, then the system is halted and an return material authorization (RMA) can then be issued for return of the failed part to the OEM. Crash events and self-recovery attempts preferably are logged by the ISSR system and provided to the OEM.


