Server Power Management During Supply Failure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern servers with multiple power supplies face challenges in maintaining operation when one power supply fails, as the total power consumption exceeds the capacity of n−1 power supplies in redundant configurations, leading to potential server downtime.
Innovation Solution
A method that identifies non-essential hardware components and selectively removes power from them to ensure the central processing unit and memory device can continue running by dynamically adjusting power consumption levels and migrating jobs to maintain server functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple power supplies are used to provide redundancy, then system reliability is improved, but the power consumption exceeds the capacity of remaining power supplies when one fails
Solution Approach 1:
The system dynamically adjusts the operational state of hardware components based on power supply status. When a power supply fails, the workload manager identifies and powers down non-essential hardware components, transitioning the system from a static power consumption state to a dynamic adaptive state that matches available power capacity.
Solution Approach 2:
The invention changes the power consumption parameter of hardware components by transitioning them between powered and powered-down states. This parameter change allows the total power consumption to be reduced to match the capacity of remaining power supplies, resolving the contradiction between reliability and energy use.
2Productivity
If all hardware components are powered to run all jobs, then productivity is improved, but system reliability deteriorates when power supply capacity is exceeded
Solution Approach 1:
The workload manager segments hardware components into essential and non-essential categories, and segments jobs into those that can continue and those that must be migrated. This segmentation allows critical functions to maintain productivity while non-critical functions are powered down to ensure system reliability.
Solution Approach 2:
The system applies partial action by powering down only the non-essential hardware components rather than all components. This selective approach maintains sufficient productivity for essential jobs while reducing total power consumption to match available power supply capacity.
3Use of energy by moving object
If non-essential hardware components are powered down, then power consumption is reduced, but device complexity increases due to component management
Solution Approach 1:
The workload manager automatically performs the complex task of identifying, selecting, and powering down non-essential hardware components without manual intervention. This self-service approach handles the device complexity internally, allowing the system to reduce power consumption while maintaining ease of operation for the user.
4Reliability
If critical jobs are prioritized on remaining power supplies, then reliability is improved, but productivity decreases for non-critical jobs
Solution Approach 1:
The workload manager acts as an intermediary that coordinates between power supply status, hardware component states, and job execution. It mediates the conflict between reliability and productivity by selectively allocating remaining power capacity to critical jobs while managing the power-down of non-essential components.
Data Source
AI summary
A method includes supplying power to a physical server from a plurality of power supplies, wherein operation of all hardware components of the server requires more power than any one of the power supplies can provide. A plurality of jobs are run on the server while the plurality of power supplies are supplying power to the physical server. The method further comprises identifying an amount of power required by each of the components, and identifying one or more components that are not required by one or more of the jobs that are running on the server. The method detects a loss of power from one of the power supplies and then selectively removes power from hardware components identified as not required so that at least a central processing unit and a memory device can continue running at least one job using power available from the operational power supplies.


