Co-Processing Unit Fault Detection and Context Restoration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current solutions for detecting fault conditions in computing systems, particularly in GPUs, result in loss of user context information, requiring users to restart applications and leading to dissatisfaction due to data loss and disruption in service.
Innovation Solution
A co-processing unit is used to detect fault conditions and restore the processing unit using stored pre-fault user context information, allowing seamless recovery without rebooting drivers or applications, thus maintaining system functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the GPU is rebooted to restore functionality after a fault condition, then the processing unit can resume operation, but user context information is lost requiring application restart
Solution Approach 1:
The patent applies preliminary action by saving user context information to a file in the file system before a fault condition occurs. When the GPU faults and is rebooted, the preserved context information is restored from the file, eliminating the need to restart applications. This proactive preservation of state information before failure resolves the contradiction between restoring GPU functionality and maintaining user context.
2Reliability
If the GPU driver is rebooted to restore the processing unit, then system functionality is recovered, but service disruption occurs leading to user dissatisfaction
Solution Approach 1:
By pre-saving context information to the file system before faults occur, the system enables rapid restoration without driver or application restarts. This eliminates service disruption time while maintaining functionality recovery, resolving the contradiction between reliable system recovery and continuous service operation.
3Ease of manufacture
If inadequate shielding is used in GPU design to reduce manufacturing costs, then device functionality and size are optimized, but susceptibility to discharging events increases
Solution Approach 1:
The patent converts the harmful effect of electrostatic discharge events that cause GPU faults into a beneficial system. By implementing automatic fault detection, context information preservation to the file system, and automated restoration capabilities, the system transforms unavoidable hardware vulnerabilities into an resilient computing platform that recovers automatically without user intervention, effectively neutralizing the impact of inadequate shielding.
Data Source
AI summary
A processing unit of a system detects a fault condition associated with the co-processing unit and, upon detection, restores the processing unit using stored user context information. During normal operation, user context information used to execute operation commands are stored by the co-processing unit in memory and maintained after fault detection. A fault condition is detected when at least a portion of the processing unit is rendered non-operational due to a discharging electrostatic event. Fault conditions may be detected by receiving information by the co-processing unit indicative of a fault condition, or by checking at least one memory location associated with processing unit to determine if information stored therein indicates a fault condition. The co-processing unit returns the processing unit to a known, workable state by using the stored user context information to restore the pre-fault detection state information to the memory locations associated with the processing unit.


