GPU Liquid-Assisted Air Cooling Failure Response for Workload Continuity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Failure of the liquid assisted air cooling (LAAC) system in data processing systems can lead to overheating and damage of hardware components, resulting in workload loss and delays in providing computer-implemented services.

Innovation Solution

A management controller is implemented to manage LAAC systems by identifying types of failures and initiating customized automatic failure responses, such as workload migration and redundancy mechanisms, to preserve operable conditions and minimize workload loss.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Temperature

If liquid assisted air cooling (LAAC) system is used to cool hardware components, then temperature control is improved, but system complexity increases due to potential failures

Engineering Contradiction:
Improvehardware component temperatureVSAvoidcooling system complexity
Core Design Contradiction:
TemperatureVSDevice complexity

Solution Approach 1:

The patent implements predictive monitoring of LAAC system health parameters (vibration, temperature, flow rate) to detect degradation trends before actual failure occurs. This allows proactive workload migration or redundancy activation, cushioning against the harmful effects of cooling system failure without requiring immediate system shutdown.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The patent introduces a management controller as an intermediary between the LAAC system and the hardware components. This controller monitors cooling system status and coordinates workload migration or redundancy activation, decoupling the direct dependency between cooling system failure and hardware operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If LAAC system failure is allowed to occur, then device complexity is reduced, but reliability deteriorates due to hardware damage risk

Engineering Contradiction:
Improvecooling system complexityVSAvoidhardware operation reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements preliminary actions by pre-configuring redundant hardware components and establishing workload migration pathways before LAAC system failure occurs. When degradation is detected, the system automatically activates pre-prepared redundancy or migrates workloads, ensuring continuous reliable operation without requiring complex real-time decision-making during failure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous feedback loops that monitor LAAC system health parameters and automatically trigger reliability-preserving actions. The management controller receives feedback from sensors monitoring vibration, temperature, and flow rate, and adjusts system operation accordingly to maintain hardware reliability throughout the cooling system's lifecycle.

Inventive Principle:
Principle #23Feedback

3Reliability

If automatic failure response is implemented, then workload loss is reduced, but device complexity increases

Engineering Contradiction:
Improveworkload continuityVSAvoidmanagement system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service automation where the management controller autonomously monitors LAAC system health, detects degradation patterns, and executes workload migration or redundancy activation without human intervention. This automated self-service approach reduces workload loss while managing complexity through intelligent algorithms rather than manual procedures.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent monitors changes in LAAC system parameters (vibration frequency, temperature gradients, flow rate variations) to detect degradation patterns. By tracking parameter changes over time rather than relying on binary failure states, the system can predict failures and initiate protective actions, improving workload continuity with manageable complexity.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively reduces the impact of LAAC system failures by maintaining optimal operating conditions and preserving workloads, ensuring continuous operation of graphics processing units (GPUs) despite partial cooling system failures.

Implementation Method 1

a liquid assisted air cooling (LAAC) system... pumps to circulate liquid to cool the hardware components

Methodology Applied
Scientific EffectLiquid assisted air cooling: Convection

Data Source

PatentUS12366903B2System and method of protecting workloads for GPU complex servers with liquid assisted air cooling
Publication Date: 2025.07.22 DELL PROD LP
  • US12366903B2 patent drawing
  • US12366903B2 patent drawing
  • US12366903B2 patent drawing

AI summary

Methods, systems, and devices for providing computer-implemented services are disclosed. To provide the computer-implemented services, graphics processing units cooled using liquid assisted air cooling may be utilized. In the event of failure of liquid assisted air cooling, the extent of the failure may be identified. A response based on the extent of the failure may be implemented to reduce an impact on workloads performed by the graphics processing units. By customizing the response, the workloads may be preserved even when cooling for some of the graphics processing units is unavailable.