Accelerator Module Power Cycling Without Full Node Reboot

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods require power cycling the entire node when an individual accelerator module encounters an error, affecting all CPUs and accelerator modules, leading to prolonged downtime and increased Annual Interruption Rate (AIR) for cloud tenants.

Innovation Solution

Implementing a management controller to monitor individual accelerator modules and enable autonomous power cycle control, allowing selective reboot of impacted modules without affecting others, using a multiplexer/demultiplexer to manage power supply.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If standardized power cycle approaches are used when an accelerator module encounters an error, then the error can be resolved by rebooting the node, but the entire node must be powered down and back up, causing prolonged downtime for all CPUs and accelerator modules

Engineering Contradiction:
Improveerror resolution capabilityVSAvoiddowntime for unaffected modules
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the power cycle control from node-level to accelerator module-level by introducing a management controller that can independently power cycle individual accelerator modules. The management controller receives error indications from specific accelerator modules and sends power cycle commands only to those affected modules, rather than powering down the entire node. This segmentation allows unaffected CPUs and accelerator modules to continue operating without interruption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The management controller serves as an intermediary between the accelerator modules and the power supply system. It monitors error states from accelerator modules via indication signals and controls the power supply to specific modules based on their operational status. This intermediary enables selective power cycling of only the faulty accelerator module while keeping the rest of the node operational.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the entire node is power cycled to address an accelerator module error, then the error is resolved, but multiple cloud tenants are affected by the error instead of just one, increasing the Annual Interruption Rate (AIR)

Engineering Contradiction:
Improveerror resolution capabilityVSAvoidservice availability for cloud tenants
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the power control at the accelerator module level rather than node level. When an error is detected in a specific accelerator module, the management controller sends a power cycle command only to that individual module. This segmentation ensures that only the affected module is restarted, while other accelerator modules and CPUs continue to serve their respective cloud tenants without interruption, thereby maintaining service availability and reducing AIR.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by providing differentiated power cycle control to different accelerator modules based on their individual operational status. Each accelerator module can be independently monitored and controlled, allowing the system to apply the power cycle action only where needed (the faulty module) rather than uniformly across the entire node. This localized approach preserves productivity for unaffected cloud tenants.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If PCIe interface standards are followed for accelerator modules, then compatibility and scaling are enabled, but power cycle events cannot be implemented at the functional level, preventing hot swapping of individual accelerator modules

Engineering Contradiction:
Improvecompatibility and scalingVSAvoidhot swap capability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The management controller acts as an intermediary that enables hot swap capability despite PCIe interface limitations. It monitors the operational status of accelerator modules and controls their power supply independently, allowing individual modules to be powered down and replaced without affecting other modules or requiring a full node reboot. This intermediary layer provides the functional-level power cycle control that the PCIe interface standard does not natively support.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service hot swapping by allowing individual accelerator modules to be monitored and controlled independently. When an error is detected or maintenance is needed, the management controller can autonomously power cycle the specific module without requiring manual intervention to shut down the entire node or affect other operational modules. This self-service capability improves ease of operation while maintaining PCIe compatibility.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12608330B2Individual power cycle control of accelerator modules configured on a node
Publication Date: 2026.04.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12608330B2 patent drawing
  • US12608330B2 patent drawing
  • US12608330B2 patent drawing

AI summary

Disclosed herein is a system for implementing a management controller on a node, or network server, that is dedicated to monitoring the individual health of a plurality of accelerator modules configured on the node. Based on the monitored health, the management controller is configured to implement autonomous power cycle control of individual accelerator modules. The autonomous power cycle control is implemented without violating the requirements of standards established for accelerator modules (e.g., OPEN COMPUTE PROJECT requirements, PERIPHERAL COMPONENT INTERCONNECT EXPRESS (PCIe) interface requirements).