Distributed Power Budget Controllers for Thermal Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Centralized power and thermal management in computer systems faces challenges when devices suddenly increase power consumption, leading to processing load imbalances and potential failures, particularly in distributed systems where sudden changes can exceed power limits.

Innovation Solution

A distributed system with component controllers and a service processor that computes local power budgets and detects failures, allowing for reset of faulty components without interrupting operations, and transitions to safe mode to manage power and thermal thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If centralized power management is implemented, then power and thermal management can be coordinated across devices, but processing load on the management entity increases when devices suddenly increase power consumption

Engineering Contradiction:
Improvepower management stabilityVSAvoidprocessing load on management entity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent divides the centralized power management entity into multiple distributed component controllers, each responsible for managing power consumption of specific devices. This segmentation distributes the processing load across multiple independent controllers, preventing any single entity from becoming a bottleneck when devices suddenly increase power consumption, while maintaining coordinated power management across the system.

Inventive Principle:
Principle #1Segmentation

2Productivity

If distributed control is implemented, then processing load is reduced on management entities, but system complexity increases due to multiple controllers

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidnumber of controllers
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines multiple component controllers into a unified distributed control architecture where each controller manages specific devices but they all participate in a coordinated system. The service processor merges the functions of monitoring and coordinating these distributed controllers, allowing the system to achieve both reduced processing load on individual entities and simplified overall management through centralized oversight of the distributed components.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If reset threshold is not exceeded, then faulty controller can be reset without interrupting operations, but system reliability may be compromised if failure is not properly handled

Engineering Contradiction:
Improvecontinuous operation capabilityVSAvoidfailure recovery assurance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the service processor continuously monitors component controllers and detects failures. When a failure is detected, the service processor evaluates whether the reset threshold has been exceeded and determines the appropriate recovery action. This feedback loop ensures that faulty controllers are reset appropriately without interrupting operations when safe, while maintaining system reliability through proper failure handling and threshold-based decision making.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9939867B2Handling a failure in a system with distributed control of power and thermal management
Publication Date: 2018.04.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9939867B2 patent drawing
  • US9939867B2 patent drawing
  • US9939867B2 patent drawing

AI summary

An apparatus includes a plurality of components and a plurality of component controllers. Each of the plurality of component controllers is associated with at least one component of the plurality of components. Each component controller is configured to compute a local power budget for the at least one component based, at least in part, on the power differential and the proportion of the total power consumption corresponding to the at least one component. A service processor is configured to determine failure associated with at least one component controller of the plurality of component controllers or the at least one component associated with the at least one component controller. The service processor is configured to in response to a reset threshold not being exceeded, reset the at least one component controller without interrupting operations of any components of the at least one component that have not failed.