Accelerator Module Power Cycling Without Full Node Reboot
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods require power cycling the entire node when an individual accelerator module encounters an error, affecting all CPUs and accelerator modules, leading to prolonged downtime and increased Annual Interruption Rate (AIR) for cloud tenants.
Innovation Solution
Implementing a management controller to monitor individual accelerator modules and enable autonomous power cycle control, allowing selective reboot of impacted modules without affecting others, using a multiplexer/demultiplexer to manage power supply.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If standardized power cycle approaches are used when an accelerator module encounters an error, then the error can be resolved by rebooting the node, but the entire node must be powered down and back up, causing prolonged downtime for all CPUs and accelerator modules
Solution Approach 1:
The patent segments the power cycle control from node-level to accelerator module-level by introducing a management controller that can independently power cycle individual accelerator modules. The management controller receives error indications from specific accelerator modules and sends power cycle commands only to those affected modules, rather than powering down the entire node. This segmentation allows unaffected CPUs and accelerator modules to continue operating without interruption.
Solution Approach 2:
The management controller serves as an intermediary between the accelerator modules and the power supply system. It monitors error states from accelerator modules via indication signals and controls the power supply to specific modules based on their operational status. This intermediary enables selective power cycling of only the faulty accelerator module while keeping the rest of the node operational.
2Reliability
If the entire node is power cycled to address an accelerator module error, then the error is resolved, but multiple cloud tenants are affected by the error instead of just one, increasing the Annual Interruption Rate (AIR)
Solution Approach 1:
The patent segments the power control at the accelerator module level rather than node level. When an error is detected in a specific accelerator module, the management controller sends a power cycle command only to that individual module. This segmentation ensures that only the affected module is restarted, while other accelerator modules and CPUs continue to serve their respective cloud tenants without interruption, thereby maintaining service availability and reducing AIR.
Solution Approach 2:
The patent applies local quality by providing differentiated power cycle control to different accelerator modules based on their individual operational status. Each accelerator module can be independently monitored and controlled, allowing the system to apply the power cycle action only where needed (the faulty module) rather than uniformly across the entire node. This localized approach preserves productivity for unaffected cloud tenants.
3Adaptability or versatility
If PCIe interface standards are followed for accelerator modules, then compatibility and scaling are enabled, but power cycle events cannot be implemented at the functional level, preventing hot swapping of individual accelerator modules
Solution Approach 1:
The management controller acts as an intermediary that enables hot swap capability despite PCIe interface limitations. It monitors the operational status of accelerator modules and controls their power supply independently, allowing individual modules to be powered down and replaced without affecting other modules or requiring a full node reboot. This intermediary layer provides the functional-level power cycle control that the PCIe interface standard does not natively support.
Solution Approach 2:
The system enables self-service hot swapping by allowing individual accelerator modules to be monitored and controlled independently. When an error is detected or maintenance is needed, the management controller can autonomously power cycle the specific module without requiring manual intervention to shut down the entire node or affect other operational modules. This self-service capability improves ease of operation while maintaining PCIe compatibility.
Data Source
AI summary
Disclosed herein is a system for implementing a management controller on a node, or network server, that is dedicated to monitoring the individual health of a plurality of accelerator modules configured on the node. Based on the monitored health, the management controller is configured to implement autonomous power cycle control of individual accelerator modules. The autonomous power cycle control is implemented without violating the requirements of standards established for accelerator modules (e.g., OPEN COMPUTE PROJECT requirements, PERIPHERAL COMPONENT INTERCONNECT EXPRESS (PCIe) interface requirements).


