GPU Power Management During Distributed Deep Learning Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing GPU power management methods are inefficient for distributed deep learning, particularly in data and tensor parallelism, leading to energy inefficiency without performance degradation.

Innovation Solution

Optimize GPU voltage and frequency during communication sections in distributed deep learning by identifying and controlling these parameters to minimum levels without degrading performance, using pipeline, data, or tensor parallelism.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If GPU voltage and frequency are maintained at high levels during distributed deep learning training, then training performance is preserved, but energy consumption increases

Engineering Contradiction:
Improvetraining performanceVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent dynamically adjusts GPU voltage and frequency based on real-time operation phases. During communication sections between data parallelism operations, the GPU operates at reduced voltage and frequency, while during compute-intensive training sections, it operates at full performance levels. This dynamic adaptation resolves the contradiction by matching power consumption to actual computational needs.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the voltage and frequency parameters of the GPU according to the operational phase. By identifying communication sections in the training pipeline and reducing voltage/frequency during these periods, the system achieves lower energy consumption without permanently sacrificing training performance, as full parameters are restored during compute sections.

Inventive Principle:
Principle #35Parameter changes

2Loss of energy

If GPU voltage and frequency are reduced to minimize energy consumption, then energy efficiency improves, but training performance degrades

Engineering Contradiction:
Improveenergy efficiencyVSAvoidtraining performance
Core Design Contradiction:
Loss of energyVSProductivity

Solution Approach 1:

The patent segments the distributed deep learning training process into distinct phases: compute-intensive training sections and communication sections. By applying different voltage and frequency settings to each segment, the system reduces energy consumption during communication phases without impacting overall training performance, as the compute sections maintain full performance levels.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements periodic adjustment of GPU power parameters synchronized with the training pipeline rhythm. The GPU alternates between high-performance mode during training computations and low-power mode during communication operations, creating a periodic pattern that maintains overall performance while improving energy efficiency.

Inventive Principle:
Principle #19Periodic action

3Adaptability or versatility

If GPU power management is applied to data parallelism and tensor parallelism, then energy efficiency improves across more architectures, but system complexity increases

Engineering Contradiction:
Improveapplicability to parallelism typesVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal power management mechanism that works across multiple parallelism types including data parallelism, tensor parallelism, and pipeline parallelism. The same voltage and frequency adjustment methodology applies to all these architectures, achieving broad adaptability without requiring architecture-specific implementations, thus limiting the increase in system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250264922A1Apparatus and method for GPU power management in distributed deep learning
Publication Date: 2025.08.21 ELECTRONICS & TELECOMM RES INST
  • US20250264922A1 patent drawing
  • US20250264922A1 patent drawing
  • US20250264922A1 patent drawing

AI summary

Disclosed herein is an apparatus and method for Graphics Processing Unit (GPU) power management in distributed deep learning. In the method, a training process or inference process of the distributed deep learning is performed through two or more GPUs, and the method may include identifying at least one communication section during the training process or inference process of the distributed deep learning and performing control such that the voltage and frequency of the GPU are optimized during the at least one communication section.