Co-Processing Unit Fault Detection and Context Restoration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current solutions for detecting fault conditions in computing systems, particularly in GPUs, result in loss of user context information, requiring users to restart applications and leading to dissatisfaction due to data loss and disruption in service.

Innovation Solution

A co-processing unit is used to detect fault conditions and restore the processing unit using stored pre-fault user context information, allowing seamless recovery without rebooting drivers or applications, thus maintaining system functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the GPU is rebooted to restore functionality after a fault condition, then the processing unit can resume operation, but user context information is lost requiring application restart

Engineering Contradiction:
ImproveGPU operational reliabilityVSAvoiduser context information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by saving user context information to a file in the file system before a fault condition occurs. When the GPU faults and is rebooted, the preserved context information is restored from the file, eliminating the need to restart applications. This proactive preservation of state information before failure resolves the contradiction between restoring GPU functionality and maintaining user context.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the GPU driver is rebooted to restore the processing unit, then system functionality is recovered, but service disruption occurs leading to user dissatisfaction

Engineering Contradiction:
Improvesystem functionality recoveryVSAvoidservice disruption time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

By pre-saving context information to the file system before faults occur, the system enables rapid restoration without driver or application restarts. This eliminates service disruption time while maintaining functionality recovery, resolving the contradiction between reliable system recovery and continuous service operation.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If inadequate shielding is used in GPU design to reduce manufacturing costs, then device functionality and size are optimized, but susceptibility to discharging events increases

Engineering Contradiction:
Improvemanufacturing cost and device sizeVSAvoidsusceptibility to discharging events
Core Design Contradiction:
Ease of manufactureVSObject-affected harmful factors

Solution Approach 1:

The patent converts the harmful effect of electrostatic discharge events that cause GPU faults into a beneficial system. By implementing automatic fault detection, context information preservation to the file system, and automated restoration capabilities, the system transforms unavoidable hardware vulnerabilities into an resilient computing platform that recovers automatically without user intervention, effectively neutralizing the impact of inadequate shielding.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS7702955B2Method and apparatus for detecting a fault condition and restoration thereafter using user context information
Publication Date: 2010.04.20 QUALCOMM INC
  • US7702955B2 patent drawing
  • US7702955B2 patent drawing
  • US7702955B2 patent drawing

AI summary

A processing unit of a system detects a fault condition associated with the co-processing unit and, upon detection, restores the processing unit using stored user context information. During normal operation, user context information used to execute operation commands are stored by the co-processing unit in memory and maintained after fault detection. A fault condition is detected when at least a portion of the processing unit is rendered non-operational due to a discharging electrostatic event. Fault conditions may be detected by receiving information by the co-processing unit indicative of a fault condition, or by checking at least one memory location associated with processing unit to determine if information stored therein indicates a fault condition. The co-processing unit returns the processing unit to a known, workable state by using the stored user context information to restore the pre-fault detection state information to the memory locations associated with the processing unit.