Neural Network Processor Fault Isolation for Fast Task Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional fault handling in neural network processors involves restarting the entire processor upon detection of a fault, leading to delayed processing and feedback, which impedes task execution and affects safety in systems like autonomous driving.
Innovation Solution
A method and apparatus for neural network processors that identify fault types and handle faults in specific modules using a preset regulation mode, allowing the processor to quickly recover and continue task execution without restarting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the entire neural network processor is restarted upon fault detection, then the fault is handled to restore system operation, but the processing time is substantially consumed and task execution is delayed
Solution Approach 1:
The patent segments the neural network processor into multiple independent functional modules (computing module, storage module, interface module, etc.). When a fault occurs, only the faulty module is restarted rather than the entire processor. This modular segmentation allows isolated fault handling, maintaining system operation in non-faulty modules and significantly reducing the time loss associated with full system restarts.
2Reliability
If the entire neural network processor is restarted upon fault detection, then the system is restored to normal operation, but the task execution progress is affected
Solution Approach 1:
By segmenting the processor into independent modules with separate control and state management, the patent enables selective restart of only the faulty module. This allows other modules to continue executing their tasks without interruption, maintaining overall task execution progress while restoring system reliability through targeted fault handling.
Solution Approach 2:
The patent implements preliminary fault detection mechanisms that identify faults before they cause complete system failure. By detecting faults early and initiating module-level restart procedures, the system can restore normal operation with minimal disruption to ongoing task execution, preserving productivity while ensuring reliability.
3Reliability
If the entire neural network processor is restarted upon fault detection, then the fault is resolved, but the response to external information is delayed
Solution Approach 1:
The modular architecture enables independent restart of faulty modules without affecting the response capability of other modules. External information processing can continue in non-faulty modules, maintaining system response speed while resolving faults in specific segments through targeted restart operations.
Solution Approach 2:
The patent ensures continuous useful action by maintaining operation of non-faulty modules during fault handling. While the faulty module is restarted, other modules continue processing external information and executing tasks, preventing complete system idle time and preserving overall response speed.
Data Source
AI summary
Disclosed are a fault handling method for a neural network processor, comprising: obtaining fault information of the neural network processor; determining a fault type of a faulty module in the neural network processor according to the fault information; and handling the fault in the faulty module using a preset regulation mode according to the fault type. When a fault is detected in the neural network processor, the method described above first determine the fault type of the current fault, and then select an appropriate regulation mode to handle the fault in the faulty module. This enables the neural network processor to be quickly restored to a normal working state and continue executing tasks that were interrupted by the fault, thereby improving the fault handling efficiency of the neural network processor. This ensures that the autonomous driving system can respond quickly to external information without affecting the execution progress of tasks.


