GPU Cluster Allocation Using Model-Specific Power Profiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cloud computing systems fail to optimize GPU resource allocation for machine learning tasks based on energy efficiency, leading to inefficient energy consumption due to varying energy requirements of different ML models and GPUs.

Innovation Solution

A method and apparatus for configuring a GPU cluster by measuring power consumption characteristics of each server for different models and using integer programming to assign servers such that the total power consumption is minimized, considering maximum throughput and model-specific energy consumption patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If newer GPUs are used for all tasks, then energy efficiency is improved, but device complexity and cost increase

Engineering Contradiction:
Improveenergy efficiencyVSAvoiddevice complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent applies local quality by assigning different GPU types to different ML models based on their specific energy consumption characteristics. Instead of using newer GPUs for all tasks, the system measures and analyzes power consumption for each GPU-model combination, then optimally assigns GPUs to models where they provide the best energy efficiency. This resolves the contradiction by making GPU allocation heterogeneous and model-specific rather than uniform.

Inventive Principle:
Principle #3Local quality

2Productivity

If GPU allocation is based on performance prioritization, then productivity is improved, but energy efficiency deteriorates

Engineering Contradiction:
ImproveproductivityVSAvoidenergy efficiency
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent changes the allocation parameter from performance-based to energy consumption-based. Instead of allocating GPUs according to raw performance metrics, the system measures actual power consumption for each GPU-model combination and uses this energy parameter as the primary criterion for allocation. This resolves the contradiction by shifting the optimization objective from productivity to energy efficiency while maintaining adequate performance through the three conditions.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If GPU servers are assigned without considering model-specific energy profiles, then ease of operation is improved, but energy consumption increases

Engineering Contradiction:
Improveease of operationVSAvoidenergy consumption
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The patent applies preliminary action by measuring and storing power consumption characteristics for each GPU-model combination before actual task allocation. The system pre-establishes an energy consumption profile database that captures the specific energy requirements of different ML models on different GPUs. This resolves the contradiction by performing the complex measurement and analysis work in advance, making the actual allocation process simple while achieving optimal energy efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12554555B2Method and apparatus for configuring cluster for machine learning service based on minimizing power consumption of GPU servers
Publication Date: 2026.02.17 ELECTRONICS & TELECOMM RES INST
  • US12554555B2 patent drawing
  • US12554555B2 patent drawing
  • US12554555B2 patent drawing

AI summary

Provided is a method of configuring a cluster, which is a method of assigning graphics processing unit (GPU) servers in a cloud in which a plurality of machine learning (ML) services are executed using an apparatus for configuring a cluster. The apparatus for configuring a cluster is configured to measure the power consumption characteristics of each of the GPU servers constituting the cloud for each of a plurality of different models processing the plurality of ML services and assign at least one GPU server to each of the plurality of models using power consumption characteristics of each of the GPU servers for each of the plurality of models to configure a GPU cluster for each of the plurality of models.