Method and apparatus for allocating GPU resources

Dynamic GPU resource allocation using MIG instances and GPU Direct RDMA optimizes GPU utilization and communication, addressing inefficiencies in multi-tenant environments by flexibly allocating resources and minimizing interference and bottlenecks.

WO2026116572A1PCT designated stage Publication Date: 2026-06-04ACRIIL

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ACRIIL
Filing Date
2024-12-23
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Current GPU resource allocation methods, particularly in multi-tenant environments, suffer from inefficiencies due to fixed-size MIG instances and limited data exchange technologies, leading to resource interference and communication bottlenecks, which reduce utilization and increase operating costs.

Method used

A dynamic GPU resource allocation method that divides GPUs into MIG instances of predetermined units, dynamically configures resources based on task requirements, and uses GPU Direct RDMA for efficient data exchange between instances, minimizing interference and bottlenecks.

Benefits of technology

This approach maximizes GPU resource utilization, reduces operating costs, and enhances performance by allowing flexible allocation and efficient communication, especially in distributed learning environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020942_04062026_PF_FP_ABST
    Figure KR2024020942_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A method for dynamically allocating GPU resources, according to one embodiment, may comprise the steps of: receiving a resource requirement of a job requested by a user; partitioning hardware resources of a GPU into a plurality of multi-instance GPU (MIG) instances on the basis of the received resource requirement, and configuring each MIG instance as a unit instance; dynamically generating an MIG instance of a target resource size by combining the unit instances in order to provide resources suitable for the job; allocating the generated MIG instance of the target resource size to the job; monitoring resource usage of the GPU hardware in real time during execution of the job, and suspending or reconfiguring the job when the resource usage exceeds a preset threshold; and performing dynamic allocation by repeating the foregoing steps for a plurality of jobs.
Need to check novelty before this filing date? Find Prior Art