Slice-Based Processing Unit Isolation for Permanent Safety Faults
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for addressing permanent faults in automotive systems are inefficient, often requiring redundant hardware that increases cost without improving performance, and result in prolonged system inoperability.
Innovation Solution
Implementing a slice-based architecture that allows for dynamic power collapsing of faulty slices in processing units, identifying transient or permanent faults through built-in self-test techniques, and rerouting workloads to functional slices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundant hardware is used to address permanent faults, then system reliability is improved, but device complexity and cost increase
Solution Approach 1:
The processing unit is divided into multiple scalable blocks that can be independently managed. When a permanent fault is detected in one block, only that specific block is power-collapsed while other blocks continue to operate, avoiding the need for redundant hardware across the entire system.
Solution Approach 2:
The system dynamically changes the operational state of processing blocks by power-collapsing faulty blocks while maintaining power to functional blocks. This parameter change allows the system to adapt to permanent faults without requiring redundant hardware components.
2Productivity
If the system continues operation after permanent fault, then productivity is maintained, but system safety may be compromised
Solution Approach 1:
The faulty scalable block is extracted from the operational system through power-collapsing, removing the source of permanent faults while allowing the remaining functional blocks to continue operation. This extraction maintains safety by isolating the fault while preserving productivity through continued operation of healthy blocks.
Solution Approach 2:
The power management mechanism acts as an intermediary between fault detection and system operation. It mediates by power-collapsing the faulty block to prevent fault propagation while allowing the system to continue operating with remaining blocks, thus maintaining both safety and productivity.
3Device complexity
If scalable blocks are power-collapsed to handle permanent faults, then device complexity is reduced, but loss of time occurs during fault response
Solution Approach 1:
The system performs preliminary classification of faults as permanent or transient before taking corrective action. This preliminary action enables the system to quickly determine whether power-collapsing is necessary, reducing the time lost during fault response by avoiding unnecessary operational disruptions.
4Reliability
If transient faults are re-executed after predetermined time, then reliability is improved, but loss of time increases due to waiting
Solution Approach 1:
The system implements periodic re-execution of workloads on scalable blocks after a predetermined time interval following transient faults. This periodic action allows temporary faults to self-correct while maintaining system reliability, balancing the need for fault tolerance with minimal time loss through structured retry mechanisms.
Data Source
AI summary
A method includes detecting a functional safety fault and determining whether the functional safety fault is a transient fault or a permanent fault. The method further includes identifying a scalable block of a processing unit from which the functional safety fault occurred in response to the functional safety fault being determined to be a permanent fault. The method optionally includes power collapsing the scalable block. The method also includes notifying a scheduler to prevent the scheduler from scheduling future workloads to the scalable block.


