Server Power Management via Dynamic Node Load Redistribution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High redundant power output capacity in data server systems leads to inefficient resource utilization and unstable operation when all power-supplying units are normal, and existing solutions that reduce power consumption to maintain system performance are not effective in managing power supply status and consumption dynamically.

Innovation Solution

A method for system power management that detects real-time power outputs and consumptions of multiple power-supplying units and computing nodes, calculates auxiliary power consumption, and adjusts node power consumption by redistributing power among nodes to maintain system stability and performance when a power-supplying unit malfunctions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If redundant power-supplying units are configured to ensure system reliability, then system reliability is improved, but redundant power output capacity becomes idle resource when all PSUs are normal

Engineering Contradiction:
Improvesystem reliabilityVSAvoididle power output capacity
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system dynamically adjusts the operating state of power-supplying units based on real-time conditions. When all PSUs are normal, the system operates in first mode utilizing only necessary power capacity. When a PSU fails, the system transitions to second mode activating the standby PSU. This dynamic adaptation eliminates idle power capacity while maintaining reliability through on-demand activation of redundant units.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters of power-supplying units based on system state. In normal conditions, redundant PSUs are placed in standby mode with reduced power output. Upon detecting a PSU failure, the system modifies parameters to activate the standby PSU at full power output, thereby converting idle capacity into active power supply without compromising system reliability.

Inventive Principle:
Principle #35Parameter changes

2Loss of energy

If PSUs with lower individual maximum power output are used to reduce redundant capacity, then power output efficiency is improved, but system performance is lowered rapidly when a PSU is malfunctioned and system switches to ULFM

Engineering Contradiction:
Improveredundant power output capacityVSAvoidsystem performance
Core Design Contradiction:
Loss of energyVSProductivity

Solution Approach 1:

The system performs preliminary configuration by selecting PSUs with higher individual power output capacity than traditionally used. This preliminary over-provisioning appears to create redundancy but actually provides headroom that prevents performance degradation when failures occur. The extra capacity is reserved and only activated when needed, thereby maintaining productivity during failure scenarios while still reducing overall redundant capacity compared to traditional approaches.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates beforehand cushioning by configuring PSUs with higher power output capacity than the minimum required. This creates a power buffer that absorbs the impact of PSU failures. When a PSU malfunctions, the cushioning capacity prevents the need to switch to ultra-low frequency mode, thereby maintaining system performance while still reducing excessive redundancy compared to traditional dual-PSU configurations.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Reliability

If system switches to ultra-low frequency mode to maintain operation after PSU failure, then system operation is maintained, but system performance is lowered rapidly and operation becomes unstable

Engineering Contradiction:
Improvesystem operation continuityVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts operating frequency based on power supply status. Instead of automatically switching to ultra-low frequency mode upon PSU failure, the system monitors real-time power capacity and adjusts frequency accordingly. This dynamic approach maintains system operation continuity while preserving performance by operating at appropriate frequency levels matched to available power capacity, avoiding the performance degradation associated with mandatory ULFM transitions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback control by continuously monitoring the operational status of power-supplying units and adjusting system operating parameters in response. When a PSU fails, the feedback mechanism detects the change in power capacity and triggers appropriate responses including activating standby PSUs and adjusting operating frequency. This feedback-based control prevents unstable operation by maintaining balanced power supply and demand, thereby sustaining both operation continuity and performance.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3495918B1Method for system power management and computing system thereof
Publication Date: 2024.04.10 GIGA COMPUTING TECHNOLOGY CO LTD
  • EP3495918B1 patent drawingFigure 1
  • EP3495918B1 patent drawingFigure 2
  • EP3495918B1 patent drawingFigure 3

AI summary

A method for system power management includes steps of detecting power output of plural power-supplying units (130) and power consumption of plural computing node (120), so as to indirectly obtain real-time auxiliary power consumption of an auxiliary unit (140) and continuously update maximum auxiliary power consumption; when one of the power-supplying units (130) is malfunctioned, renewing the maximum sum of the power output of the other power-supplying units (130), and applying the difference of the renewed maximum sum of the power outputs and the maximum auxiliary power consumption as a first sum of the node power consumptions of the computing nodes (120); finally, according to the first sum of the node power consumptions, cutting down the power consumption of at least one of the computing nodes (120) to a first node power consumption.