Server Node Computational Bursting via Liquid Cooling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud computing systems face increasing costs due to unused computing resources and premature server node replacement, as they strive to maintain performance guarantees and reliability, leading to inefficient resource allocation and hardware lifespan management.

Innovation Solution

A resource allocation system that implements lifetime-aware computational bursting, using an efficient liquid cooling system to optimize server node performance without degrading hardware lifetime, by selectively deploying virtual machines and managing computational bursting modes based on historical data to ensure reliable service and reduce waste.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If computational bursting is engaged to enhance virtual machine performance, then service performance is improved, but hardware lifetime is degraded

Engineering Contradiction:
Improvevirtual machine performanceVSAvoidhardware lifetime
Core Design Contradiction:
SpeedVSDuration of action of stationary object

Solution Approach 1:

The system implements periodic computational bursting by alternating between high-performance bursting modes and normal operation modes. The processor engages computational bursting for a predetermined duration to meet performance demands, then returns to normal operation to allow hardware recovery and cooling, creating a periodic cycle that balances performance enhancement with hardware preservation.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system dynamically changes operational parameters by switching between different computational bursting modes with different performance levels and hardware stress characteristics. The processor selects appropriate bursting intensities based on workload demands, adjusting the balance between performance gain and hardware wear by modifying operational parameters rather than maintaining constant high-stress operation.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If buffer capacity is maintained to ensure reliability during unexpected events, then system reliability is improved, but computing resource utilization deteriorates

Engineering Contradiction:
Improvesystem reliabilityVSAvoidcomputing resource utilization
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts buffer capacity requirements based on real-time reliability needs and workload conditions. Rather than maintaining fixed buffer capacity, the system adapts the balance between buffer reserves and active computing resources according to predicted failure risks and demand patterns, optimizing the trade-off between reliability and utilization.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms that monitor hardware health, workload patterns, and failure risks to dynamically adjust buffer capacity allocations. By continuously analyzing system state and comparing it against reliability thresholds, the system optimizes buffer maintenance levels to ensure reliability only when needed, reducing wasted capacity during normal operation.

Inventive Principle:
Principle #23Feedback

3Reliability

If server nodes are replaced early to ensure reliability, then system reliability is improved, but hardware lifespan management deteriorates

Engineering Contradiction:
Improvesystem reliabilityVSAvoidhardware lifespan
Core Design Contradiction:
ReliabilityVSDuration of action of stationary object

Solution Approach 1:

The system performs preliminary assessments of hardware health and predictive failure analysis before actual failures occur. By monitoring degradation trends and predicting potential failures in advance, the system can plan replacements or maintenance activities optimally, avoiding both premature replacement and extended operation beyond safe lifetimes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service monitoring and self-diagnosis capabilities that allow server nodes to report their own health status, degradation rates, and maintenance needs. This self-awareness enables more accurate tracking of actual hardware lifespan versus replacement timing, allowing the system to optimize the balance between reliability and hardware utilization without external intervention.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances virtual machine performance, reduces the number of unused server nodes, prolongs hardware lifespan, and minimizes costs associated with powering and replacing server nodes, while predicting computing capacity more accurately.

Implementation Method 1

hardware of the plurality of server nodes is cooled by an efficient liquid cooling system

Methodology Applied
Scientific EffectLiquid cooling: Convection

Data Source

PatentUS12135998B2Managing computational bursting on server nodes
Publication Date: 2024.11.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12135998B2 patent drawing
  • US12135998B2 patent drawing
  • US12135998B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and computer readable media for optimizing resource allocation on a computing zone based on bursting data and in accordance with an allocation policy. For example, systems disclosed herein may implement a lifetime-aware allocation of computing resources to facilitate computational bursting on respective server nodes to prioritize different aspects of performance on the server nodes of a cloud computing system. The systems described herein provide a number of benefits, including, by way of example, facilitating an allocation policy that enables more densely packing virtual machines on respective server nodes, boosting individual virtual machine performance, and reducing a buffer of empty nodes that need to be maintained on one or more node clusters. This may include more densely packing virtual machines on respective server nodes.