A GPU computing instance allocation method in a cloud server cluster

CN122547549BActive Publication Date: 2026-09-08KOLUDEO (SHANDONG) ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611015544.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-08
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

GPU的内存块数量、内存块容量和计算核心数量设计是固定的,且在GPU实例化过程中,计算实例所占用内存需要按块对齐,这就导致了内存块和计算核心所支持的计算实例模式有限

Benefits of technology

本发明在选择内存分配方案时,优先选取配置后能支持最多实例模式的候选内存分配方案,确保内存块分配后仍能保留足够连续空间以支持后续实例,为计算实例的分配引入了前瞻性评估机制,当需要在多个候选内存分配方案中选择时,不是简单地选择第一个可用的方案,而是评估每个候选内存分配方案对后续分配能力的影响。这使得分配决策不仅考虑当前实例的满足,还考虑对未来实例的支持能力,从根源上减少内存被分割为零散小块的情况,维持内存连续性,显著减少GPU内存碎片化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547549B_ABST
    Figure CN122547549B_ABST
Patent Text Reader

Abstract

The application discloses a GPU computing instance allocation method in a cloud server cluster. The method comprises the following steps: acquiring configuration information of GPUs in the cluster and supported instance modes; constructing a feasible starting memory block set of any instance mode in advance; when a target computing instance is allocated to a selected GPU, acquiring an instance mode of the target computing instance and a state of the selected GPU; traversing all possible starting memory blocks to acquire candidate memory allocation schemes; according to memory distribution situations after application of each candidate memory allocation scheme, counting the number of instance modes that can be satisfied; and selecting a candidate memory allocation scheme that can satisfy the most instance modes for allocation. The application further divides instance modes into a first type of instance mode and a second type of instance mode that require complete GPUs, and respectively establishes pre-allocation GPU groups for management. The application can significantly reduce GPU memory fragmentation and improve GPU resource utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud computing resource scheduling, and in particular to a method for allocating GPU computing instances in a cloud server cluster. Background Technology

[0002] The rapid growth of machine learning, data analytics, and high-performance computing has made GPUs a core computing resource for cloud computing services. While GPUs offer superior parallel processing capabilities, significantly accelerating computationally intensive tasks such as deep learning training and scientific computing, they also present challenges in scheduling and resource sharing.

[0003] To fully utilize GPUs, existing technologies divide a single GPU into multiple isolated instances, each with dedicated memory and computing cores, enabling multiple tasks to run in parallel on the same GPU. The number of memory blocks, memory block capacity, and the number of computing cores within a GPU are designed to be fixed. Furthermore, during GPU instantiation, the memory occupied by each computing instance needs to be aligned by block, which limits the computing instance modes supported by memory blocks and computing cores. Inappropriate allocation of computing instance resources within a GPU can lead to GPU memory fragmentation. Further, within a cluster, the migration of virtual machines that allocate tasks to GPUs, resulting in frequent creation and destruction of GPU computing instances, exacerbates GPU memory fragmentation, meaning that parts of a single GPU's memory become unusable or inefficiently allocated. In cloud computing clusters with large-scale GPUs, reduced utilization of a single GPU means an increase in the number of GPUs required to execute the same amount of tasks, or a reduction in the maximum number of tasks supported by the limited number of GPUs in the cloud computing cluster. Summary of the Invention

[0004] To address the aforementioned technical problems, or at least partially address them, and to mitigate the aforementioned drawbacks, this invention provides a method for allocating GPU computing instances in a cloud server cluster.

[0005] In a first aspect, the present invention provides a method for allocating GPU computing instances in a cloud server cluster, comprising: Obtain the configuration information of the GPUs in the cluster's GPU resource pool and the instance modes of the computing instances supported by the GPUs; Based on the GPU's configuration information and the instance modes of the computing instances supported by the GPU, a set of feasible starting memory blocks for any of the instance modes is pre-constructed for the GPU. When assigning a target computing instance to a selected GPU, obtain the instance mode of the target computing instance and the state of the selected GPU; Perform the following procedure to set the memory allocation for the target compute instance on the selected GPU: The algorithm iterates through all possible starting memory blocks of the instance modes of the target computing instance. For the selected GPU, it checks whether contiguous memory blocks starting from the starting memory block are available for the target computing instance, and obtains all available contiguous memory blocks as candidate memory allocation schemes. For each candidate memory allocation scheme, it obtains the GPU memory block distribution assuming the application of the candidate memory allocation scheme. Based on the memory block distribution after applying the candidate memory allocation scheme, it iterates through all feasible starting memory blocks of all instance modes supported by the GPU, and determines whether contiguous memory blocks starting from each feasible starting memory block that are larger than the memory capacity required by each instance mode are available. Based on the availability of contiguous memory blocks, it counts the number of instance modes that the GPU can satisfy after applying each candidate memory allocation scheme. The algorithm selects the candidate memory allocation scheme that maximizes the number of satisfied instance modes or the distribution-weighted scheme that maximizes the number of satisfied instance modes as the final memory allocation scheme, and then configures the target computing instance.

[0006] Furthermore, a feasible set of starting memory blocks is pre-built for the GPU, specifically including: For any instance mode, determine all possible starting memory block indices for that instance mode based on the number of compute cores and memory blocks in the GPU configuration information; For each possible starting memory block index, determine whether a contiguous memory block starting from that index can meet the memory capacity requirements of this instance mode; Add the index of the starting memory block that meets the requirements to the set of feasible starting memory blocks for this instance pattern.

[0007] Furthermore, based on the availability of contiguous memory blocks, the number of instance modes that the GPU can satisfy after applying each candidate memory allocation scheme specifically includes: The current memory block distribution status of the GPU is obtained from the GPU's state, wherein the memory block distribution status records the occupancy of each memory block; For any selected candidate memory allocation scheme, assume the distribution of GPU memory blocks after applying the candidate memory allocation scheme; Traverse all feasible starting memory blocks for all instance modes supported by the GPU. For each possible starting memory block, starting from that starting memory block, detect the number of consecutive free memory blocks. Determine whether the space capacity of the consecutive free memory blocks is greater than the memory capacity required by the target computing instance. If so, determine that consecutive memory blocks starting from that starting memory block that are greater than the memory capacity required by each instance mode are available. Then, the GPU supports the corresponding instance mode after applying the candidate memory allocation scheme. Count the total number of instance modes supported by the GPU after applying the candidate memory allocation scheme. Traverse the candidate memory allocation schemes to obtain the total number of instance modes supported by the GPU after applying each candidate memory allocation scheme.

[0008] Furthermore, the compute instance modes are divided into a first type of instance mode and a second type of instance mode. The first type of instance mode requires a complete GPU resource for execution, while the second type of instance mode, after being allocated to a GPU, requires the GPU to support at least one minimum compute instance. The GPU resource pool in the cluster is further divided into a first type of instance mode pre-allocated GPU group and a second type of instance mode pre-allocated GPU group. The first type of instance mode pre-allocated GPU group is specifically used to execute compute instances of the first type of instance mode, and the second type of instance mode pre-allocated GPU group is specifically used to execute compute instances of the second type of instance mode.

[0009] Furthermore, based on the proportion of computing instances in the first type of instance mode, a corresponding number of GPUs are selected from the cluster's GPU resource pool to form a pre-allocated GPU group for the first type of instance mode. It is ensured that the selected GPUs come from at least a set number of servers to achieve distributed deployment. The computing instances of the first type of instance mode are then allocated to the GPUs in the pre-allocated GPU group for the first type of instance mode. Select a set number of GPUs from the remaining GPUs in the cluster's GPU resource pool to form a pre-allocated GPU group for the second type of instance mode. Ensure that the selected GPUs come from at least a set number of servers to achieve distributed deployment. Assign the computing instances of the second type of instance mode to the GPUs in the pre-allocated GPU group for the second type of instance mode.

[0010] Furthermore, the first type of instance mode pre-allocated GPU group supports dynamic adjustment, including: when the pre-allocated GPU group of the first type of computing instance mode cannot meet the current computing instance's running needs due to fluctuations in the instance mode proportion, within the maximum size limit of the pre-allocated GPU group of the first type of instance mode, selecting new empty GPUs from the cluster GPU resource pool and adding them to the pre-allocated GPU group of the first type of instance mode, and allocating the target computing instance of the first type of instance mode to the newly added GPU; detecting the number of idle GPUs in the pre-allocated GPU group of the first type of instance mode, and when the number of idle GPUs in the pre-allocated GPU group of the first type of instance mode exceeds a set threshold and reaches a set time, reclaiming the idle GPUs to the GPU resource pool, or when the number of idle GPUs in the pre-allocated GPU group of the first type of instance mode does not exceed the set threshold but reaches the reclamation period, reclaiming the idle GPUs to the GPU resource pool.

[0011] Furthermore, the second type of instance mode's pre-allocated GPU groups support dynamic adjustment based on task requirements, including: Based on the instance mode of the compute instance, detect whether there is a GPU that supports the operation in the pre-allocated GPU group of the second type of instance mode; If it exists, the size of the pre-allocated GPU group for the second type of instance mode remains unchanged, and the target computing instance is allocated to the first detectable runnable GPU. Alternatively, all candidate GPUs that support the running of computing instances are obtained. For each candidate GPU, the number of distributed weighted instance modes satisfied after applying the target computing instance is calculated, and the candidate GPU that maximizes the number of distributed weighted instance modes satisfied by the second type of instance mode is selected as the final allocation target. If it does not exist, a new empty GPU is selected from the cluster GPU resource pool and added to the pre-allocated GPU group of the second type of instance mode, and the target computing instance is assigned to the new empty GPU. If the cluster GPU resource pool is empty, then wait. Based on resource differences or a set time period, defragmentation is performed to reclaim GPUs from the pre-allocated GPU groups in the second instance mode to the GPU resource pool.

[0012] Furthermore, defragmentation based on resource differences or a set time period to reclaim GPUs from the pre-allocated GPU group of the second-type instance mode to the GPU resource pool includes: detecting the GPU resource size of the pre-allocated GPU group of the second-type instance mode and the GPU resource difference required by all the second-type instance mode computing instances it executes; if the resource difference exceeds a set threshold and reaches a set duration, or if it does not exceed the set threshold but reaches the defragmentation period, then defragmentation is performed on the pre-allocated GPU group of the second-type instance mode. During the defragmentation process, the set of GPUs belonging to the pre-allocated GPU group of the second-type instance mode on any server in the cluster is determined, and it is checked whether the GPUs in the GPU set are allowed to migrate computing instances to obtain idle GPUs. If so, local GPU migration of computing instances is performed, tasks are redirected, and idle GPUs are removed from the pre-allocated GPU group of the second-type instance mode.

[0013] Furthermore, during the defragmentation process of GPUs in the pre-allocated GPU group of the second type of instance mode, on each server local, the current free memory blocks and occupied memory blocks of each GPU are counted, and the weighted number of instance modes supported by the free memory blocks is counted. The ratio of free memory blocks to occupied memory blocks multiplied by the weighted number is used as the evaluation value for whether the GPU should perform compute instance migration. GPUs that perform compute instance migration are selected according to the size of the evaluation value.

[0014] Secondly, the present invention provides a GPU computing instance allocation device in a cloud server cluster, comprising: at least one processing unit, wherein the processing unit is connected to a storage unit via a bus unit, wherein the storage unit stores a computer program, and the processing unit implements the GPU computing instance allocation method in the cloud server cluster by running the computer program stored in the storage unit.

[0015] The technical solutions provided in the embodiments of the present invention have the following advantages compared with the prior art: This invention prioritizes candidate memory allocation schemes that can support the most instance modes after configuration when selecting memory allocation schemes. This ensures that sufficient contiguous space is reserved after memory block allocation to support subsequent instances. It introduces a forward-looking evaluation mechanism for compute instance allocation. When choosing from multiple candidate memory allocation schemes, it does not simply select the first available scheme, but evaluates the impact of each candidate scheme on subsequent allocation capabilities. This makes allocation decisions consider not only the needs of current instances but also the support capabilities for future instances, fundamentally reducing the fragmentation of memory into small blocks, maintaining memory continuity, and significantly reducing GPU memory fragmentation.

[0016] This invention predefines feasible starting memory blocks, significantly reducing the search range during allocation, lowering computational overhead, and accelerating instance deployment.

[0017] This invention distinguishes between a first-type instance mode requiring a full GPU and a second-type instance mode that can share a GPU. It designs pre-allocation groups for each type of instance mode, along with dynamic adjustment and recycling mechanisms for each GPU group. This reduces large-scale fragmentation caused by mixed allocation of first-type and second-type instance modes, maximizing GPU resource utilization. Furthermore, this invention dynamically adapts each GPU group to fluctuating demand: the pre-allocated GPU groups for the first-type instance mode initially allocate distributed GPU resources based on the proportion of first-type instance modes. When demand increases, empty GPUs are added within the maximum capacity; when idle, they are promptly recycled back to the GPU resource pool, preventing pre-allocated groups from occupying idle resources for extended periods.

[0018] The second-type instance mode pre-allocated GPU group adds new empty GPUs only when the existing GPUs cannot meet the computing needs of the second-type instance mode. At the same time, it releases inefficiently used GPUs through local instance migration, ensuring that resources are tilted towards high-demand scenarios and reducing waste.

[0019] This invention prioritizes the allocation of GPUs and memory schemes that support the largest number of modes after allocation, so that resource configuration is highly matched with actual task requirements and the probability of subsequent instance allocation is increased.

[0020] Both the first and second instance modes of this invention employ a distributed design for pre-allocated GPU groups, with the selected GPUs coming from at least multiple servers, thus avoiding the risk of single point of failure. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 The flowchart illustrates a method for allocating GPU computing instances in a cloud server cluster, provided in an embodiment of the present invention, to optimize memory allocation of computing instances in the GPU and improve fragmentation of computing instance allocation. Figure 2 In a method for allocating GPU computing instances in a cloud server cluster provided by an embodiment of the present invention, computing instances are classified and the GPUs in the cluster GPU resource pool are grouped according to the classification of computing instances to optimize the flowchart of cluster processing computing tasks. Figure 3 This is a schematic diagram illustrating the grouping of GPU resource pools in a cluster according to an embodiment of the present invention; Figure 4 A flowchart for fragmentation and organization provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a GPU computing instance allocation device in a cloud server cluster provided in an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0026] Example 1 like Figure 1As shown, the present invention provides a method for allocating GPU computing instances in a cloud server cluster, including optimizing the memory allocation of computing instances in the GPU to improve the fragmentation of computing instance allocation, including: S1: Obtain the configuration information of the GPUs in the cluster's GPU resource pool and the instance mode of the computing instances supported by the GPUs.

[0027] The GPU configuration information includes the number of GPU computing cores and memory blocks. Each computing core can independently execute computing tasks, and the computing core is the smallest unit of computing resources for the GPU. The GPU's video memory is divided into several fixed memory blocks, each with the same capacity. The storage capacity of each memory block is a fixed parameter of the GPU hardware. It is assumed that any computing core can access any memory block. The configuration of any GPU in the cluster is represented as a tuple. ,in: This represents the total number of GPU computing cores, and is a positive integer. The total number of memory blocks in the GPU, a positive integer; The capacity of a single memory block; To calculate the core state vector, Indicates the first One computing core is idle. This indicates that the i-th computing core is already in use; For memory block state vectors, Indicates the first One memory block is free. This indicates that the space is already occupied.

[0028] An instance mode refers to the GPU resource requirements of a type of compute instance, including the number of compute cores and memory space requirements. The memory space requirement for an instance mode is calculated by dividing the memory space requirement by the capacity of a single memory block and rounding up to the nearest integer. During GPU instantiation, the memory occupied by a compute instance needs to be aligned to memory blocks. This means that the memory allocation of an instance must start from the beginning of a memory block, and the memory capacity required by the instance must be an integer multiple of the memory block capacity. When the memory space required for an instance mode is greater than the capacity of a single memory block, for the current GPU memory distribution to support this instance mode, there must be a contiguous free memory block in the GPU's current memory, and the capacity of this contiguous free memory block must be greater than the memory space required by the instance mode. The instance mode of any compute instance can be represented as... ,in, The number of compute cores required for this instance mode meets the following requirements. ; The number of memory blocks required for this mode, satisfying .

[0029] For a GPU with M compute cores and N memory blocks, where any compute core can access any memory block, the maximum instance mode requires M compute cores and N times the memory block capacity in memory space; the minimum instance mode requires 1 compute core and 1 memory block capacity in memory space; the range of instance modes supported by the GPU is between the maximum and minimum modes.

[0030] S2: Based on the GPU's configuration information and the instance modes of the computing instances supported by the GPU, pre-build a set of feasible starting memory blocks for any of the instance modes for the GPU.

[0031] For any instance mode, based on the number of GPU cores and memory blocks, all possible starting memory block indices for that instance mode are determined, and those that meet the requirements are added to the feasible starting memory block set for that instance mode. Since the number of instance modes supported by a GPU is limited, and the required memory space and number of computing cores for each instance mode further restrict the resources that can be allocated to it; for example, for a GPU with 6 memory blocks and 4 computing cores, when running a computing instance requiring 5 times the memory block unit capacity, the instance can only be allocated starting from the first or second memory block. Therefore, for any instance mode, there are feasible starting memory blocks, and the feasible starting block set can be reused once constructed. The advantage is that when allocating computing instances of any mode, allocation is only performed within the feasible domain defined by the feasible starting memory blocks, avoiding a feasibility search and allocation across the entire memory space.

[0032] For any instance mode, based on the number of computing cores and memory blocks in the GPU configuration information, determine all possible starting memory block indices for that instance mode. If instance mode p requires n memory blocks, the possible memory block indices range from 1 to N-n+1. For each possible starting memory block index, determine whether a contiguous block of memory starting from that index can meet the memory capacity requirements of the instance mode. Add the starting memory block indices that meet the requirements to the feasible starting memory block set of the instance mode. The feasible starting memory block set of instance mode p is then represented as: .

[0033] Where j represents the possible starting memory block index, and k represents the memory block index within the range of memory block data required by the instance mode, starting from the possible starting memory block.

[0034] S3: When allocating the target computing instance to the selected GPU, obtain the instance mode of the target computing instance and the state of the selected GPU. The GPU state includes the currently remaining computing cores and the current memory block distribution, i.e. .

[0035] S4: Perform the following procedure to set the memory allocation for the target compute instance on the selected GPU: The algorithm iterates through all possible starting memory blocks for the instance modes of the target computing instance. For the selected GPU, it checks whether contiguous memory blocks starting from the starting memory block are available for the target computing instance, and obtains all available contiguous memory blocks as candidate memory allocation schemes. For each candidate memory allocation scheme, it obtains the GPU memory block distribution assuming the application of the candidate memory allocation scheme. Based on the memory block distribution after applying the candidate memory allocation scheme, it iterates through all feasible starting memory blocks for all instance modes supported by the GPU, and determines whether contiguous memory blocks starting from each feasible starting memory block that are larger than the memory capacity required by each instance mode are available. Based on the availability of contiguous memory blocks, it counts the number of instance modes that the GPU can satisfy after applying each candidate memory allocation scheme. The candidate memory allocation scheme that satisfies the largest number of instance modes is selected as the final memory allocation scheme, and the target computing instance is configured, allocating the data required for the target computing instance to the corresponding memory blocks. If multiple schemes satisfying the largest number of instance modes exist, the first one is selected; if no candidate memory allocation scheme exists, the allocation fails.

[0036] After selecting the optimal candidate memory allocation scheme as the final memory allocation scheme, the GPU's memory block distribution status is updated.

[0037] Specifically, based on the availability of contiguous memory blocks, the number of instance modes that the GPU can satisfy after applying each candidate memory allocation scheme includes: The current memory block distribution status of the GPU is obtained from the GPU's state, wherein the memory block distribution status records the occupancy of each memory block; For any selected candidate memory allocation scheme, assume the distribution of GPU memory blocks after applying the candidate memory allocation scheme; Traverse all feasible starting memory blocks for all instance modes supported by the GPU. For each possible starting memory block, starting from that starting memory block, detect the number of consecutive free memory blocks. Determine whether the space capacity of the consecutive free memory blocks is greater than the memory capacity required by the target computing instance. If so, determine that consecutive memory blocks starting from that starting memory block that are greater than the memory capacity required by each instance mode are available. Then, the GPU supports the corresponding instance mode after applying the candidate memory allocation scheme. Count the total number of instance modes supported by the GPU after applying the candidate memory allocation scheme. Traverse the candidate memory allocation schemes to obtain the total number of instance modes supported by the GPU after applying each candidate memory allocation scheme.

[0038] Through the above process, the memory distribution of the GPU after configuring the target computing instance can support more instance modes, thereby improving the problem of GPU memory fragmentation.

[0039] As a preferred approach, instead of selecting the candidate memory allocation scheme that maximizes the number of instance modes satisfied, the final memory allocation scheme is selected. The method is adjusted so that the candidate memory allocation scheme that maximizes the number of instance modes satisfied by the distributed weighting is selected as the final memory allocation scheme. The definition of the number of instance modes satisfied by the distributed weighting is: ; in: For the first The weight of each instance pattern in historical statistics or predictions satisfies ; The total number of instance modes supported by the GPU; This is an indicator function that shows the GPU's memory distribution after executing candidate memory allocation schemes. Able to support the i-th instance mode p i The value is 1 when the condition is met and 0 otherwise; the number of instance patterns satisfied by the weighted distribution. The physical meaning is: under a given memory distribution Below, the weighted expected value of the instance modes that the GPU can support. (Selection to make...) The largest candidate memory allocation scheme is essentially one that makes the allocated memory configuration more likely to be satisfied when future requests arrive.

[0040] Example 2 A method for allocating GPU computing instances in a cloud server cluster further includes classifying computing instances and grouping GPUs in the cluster's GPU resource pool according to these classifications to optimize cluster processing of computing tasks. Figure 2 and Figure 3 As shown, the specific process of Example 2 includes: Compute instance modes are divided into two types: Type 1 and Type 2. Type 1 instance modes require a full GPU resource for execution; once a compute instance is assigned to a Type 1 instance, the GPU cannot be used for other compute instances. For example, one Type 1 instance mode requires all GPU cores, another requires all GPU memory blocks, and yet another requires both GPU cores and memory blocks. For Type 2 instance modes, after being assigned to a GPU, the GPU must also support at least one minimum compute instance.

[0041] The GPU resource pool in the cluster is divided into two categories based on the GPU adaptation computing instance modes: a first type of instance mode pre-allocated GPU group and a second type of instance mode pre-allocated GPU group. The first type of instance mode pre-allocated GPU group is specifically used to execute computing instances of the first type of instance mode, and the second type of instance mode pre-allocated GPU group is specifically used to execute computing instances of the second type of instance mode.

[0042] If a first-type instance mode and a second-type instance mode, requiring full GPU resource execution, are allocated to the GPU in a mixed manner, first-type instance mode compute instances must wait until the GPU is completely idle before being allocated. The waiting time for first-type instance mode compute instances depends on the compute instance with the longest execution time among those ahead of it in the queue. Other compute instances ahead of it cannot be allocated after completing their operations, often resulting in a large amount of fragmented idle compute resources. This invention separates the first-type instance mode and the second-type instance mode, reducing the fragmented compute resources caused by first-type instance mode compute instances during the waiting process.

[0043] During cluster operation, the distribution of computational instances for processed tasks is generally dynamic and uneven. For example, 25% of computational instances might utilize two cores and two memory blocks, 50% might utilize one core and one memory block, and 3% might utilize M cores and N memory blocks, and so on. Initially, based on the proportion of computational instances in the first instance mode, a corresponding number of GPUs are selected from the cluster's GPU resource pool to form a pre-allocated GPU group for the first instance mode. This ensures that the selected GPUs come from at least a set number of servers, achieving distributed deployment. During subsequent task processing, computational instances in the first instance mode are allocated to GPUs in the pre-allocated GPU group for the first instance mode. A set number of GPUs are then selected from the remaining GPUs in the cluster's GPU resource pool to form a pre-allocated GPU group for the second instance mode. During subsequent task processing, computational instances in the second instance mode are allocated to GPUs in the pre-allocated GPU group for the second instance mode, and the allocation of computational instances within the GPUs is optimized based on the instance mode distribution.

[0044] During subsequent operation, the pre-allocated GPU group for the first type of instance mode can be dynamically adjusted according to the task situation, including: when the pre-allocated GPU group for the first type of computing instance mode cannot meet the current computing instance's running needs due to fluctuations in the instance mode's proportion, within the maximum size limit of the pre-allocated GPU group for the first type of instance mode, a new empty GPU is selected from the cluster GPU resource pool and added to the pre-allocated GPU group for the first type of instance mode, and the target computing instance of the first type of instance mode is allocated to the newly added GPU; the number of idle GPUs in the pre-allocated GPU group for the first type of instance mode is detected, and when the number of idle GPUs in the pre-allocated GPU group for the first type of instance mode exceeds a set threshold and reaches a set time, the idle GPUs are recycled to the GPU resource pool, or when the number of idle GPUs in the pre-allocated GPU group for the first type of instance mode does not exceed the set threshold but reaches the recycling cycle, the idle GPUs are recycled to the GPU resource pool.

[0045] During subsequent operation, the pre-allocated GPU groups in the second instance mode also support dynamic adjustment based on task requirements, including: When assigning a compute instance of the second instance mode to a pre-allocated GPU group of the second instance mode, the system checks whether there is a GPU in the pre-allocated GPU group of the second instance mode that can support the operation, based on the instance mode of the compute instance. If it exists, the size of the pre-allocated GPU group for the second type of instance mode remains unchanged, and the target computing instance is allocated to the first detectable runnable GPU. Alternatively, all candidate GPUs that support the running of computing instances are obtained. For each candidate GPU, the number of distributed weighted instances that meet the requirements after applying the target computing instance is calculated, and the candidate GPU with the largest number of distributed weighted instances that meet the requirements of the second type of instance mode is selected as the final allocation target. If it does not exist, a new empty GPU is selected from the cluster GPU resource pool and added to the pre-allocated GPU group of the second type of instance mode, and the target computing instance is assigned to the new empty GPU. If the cluster GPU resource pool is empty, then wait.

[0046] Based on resource differences or a set time period, defragmentation is performed to reclaim GPUs from the pre-allocated GPU groups in the second instance mode to the GPU resource pool: The system detects the difference between the GPU resource size in the pre-allocated GPU group for the second-type instance mode and the GPU resource requirements of all the compute instances in the second-type instance mode it executes. If the resource difference exceeds a set threshold and reaches a set duration, or if it does not exceed the set threshold but reaches the defragmentation cycle, then the pre-allocated GPU group for the second-type instance mode is defragmented. During the defragmentation process, the system determines the set of GPUs belonging to the pre-allocated GPU group for the second-type instance mode on any server in the cluster, and checks whether the GPUs in the GPU set are allowed to migrate compute instances to obtain idle GPUs. If so, the system performs local GPU migration for compute instances, redirects tasks, and removes the idle GPUs from the pre-allocated GPU group for the second-type instance mode.

[0047] The dynamic adjustment strategy for pre-allocating GPU groups in the second type of instance mode enables the memory distribution of GPUs to support more instance modes that conform to the second type of instance distribution after the target computing instance is configured, thereby improving the problem of GPU memory fragmentation in the cluster.

[0048] like Figure 4 As shown, the defragmentation process includes: During defragmentation, the set of GPUs belonging to the pre-allocated GPU group of the second type of instance mode on any server in the cluster is identified. Defragmentation is then performed on these local GPU sets. The benefits of restricting migration to local servers for defragmentation include: reduced latency from cross-server operations, lower probability of failure during cross-server migrations, ensuring business continuity through local task redirection, low migration costs, and minimal impact on users.

[0049] The system checks whether GPUs in the GPU set are allowed to perform compute instance migration. Specifically, on each server, the system counts the current free and occupied memory blocks of each GPU, calculates the number of satisfied instance modes using a weighted distribution, and uses the ratio of free to occupied memory blocks multiplied by the number of satisfied instance modes as the evaluation value for whether a GPU should perform compute instance migration. GPUs are then selected for compute instance migration based on the magnitude of this evaluation value. The evaluation value is: ; in, The amount of free memory blocks. To conserve memory blocks, idle GPUs do not perform migrations, therefore Not zero.

[0050] Perform a local migration. For the selected GPU, perform the following operations: migrate the compute instance from the selected GPU to the target GPU with the goal of minimizing migration costs, redirect relevant tasks to the new GPU, and put the selected GPU into an idle state.

[0051] In practice, the choice of migration path affects migration costs and business continuity. This application defines a migration cost function: ; in: The topological distance between GPU x and y is defined as the reciprocal of the network hop count or communication bandwidth. For migration size The time required for an instance; The potential impact of the migration on service level agreements includes increased latency and service downtime.

[0052] GPUs that become idle after migration are removed from the pre-allocated GPU group in the second instance mode and returned to the cluster GPU resource pool.

[0053] In real-world cloud computing workloads, allocation requests often exhibit bursty characteristics. Scenarios such as batch scheduling of deep learning training tasks and batch submission of scientific computing jobs generate a large number of requests within a short period. Processing these requests one by one is inefficient in two ways: multiple requests with the same instance mode will trigger the same evaluation logic, resulting in wasted computing resources; if a locking mechanism is used to protect the GPU state, frequent locking and unlocking operations will affect throughput.

[0054] This solution employs a batch scheduling mechanism triggered by both time windows and thresholds. Batch scheduling is triggered when any of the following conditions are met: the number of requests accumulated within the time window reaches the maximum number of requests; the earliest request's waiting time in the buffer exceeds the width of the time window; and the buffer's data structure uses a priority queue, sorted by request priority and arrival time. Priority calculation comprehensively considers factors such as the request's service level agreement requirements, business weight, and waiting time.

[0055] Example 3 like Figure 5 As shown, this embodiment of the invention provides a GPU computing instance allocation device in a cloud server cluster, comprising: at least one processing unit, the processing unit being connected to a storage unit via a bus unit, the storage unit serving as a computer-readable storage medium, and being used to store software programs, computer-executable programs, and modules, such as the software program, computer-executable program, and module corresponding to the GPU computing instance allocation method in a cloud server cluster according to this embodiment of the invention. The processing unit implements the aforementioned GPU computing instance allocation method in a cloud server cluster by running the software program, computer-executable program, and module stored in the storage unit, including: Of course, the computer program stored in the storage unit of the GPU computing instance allocation device in the cloud server cluster provided in the embodiments of the present invention is not limited to the method operation described above, and can also execute related operations in the GPU computing instance allocation method in the cloud server cluster provided in any embodiment of the present invention.

[0056] Example 4 This invention provides a computer-readable storage medium that stores a computer program. When the computer program is executed, it implements the GPU computing instance allocation method in the cloud server cluster.

[0057] In the embodiments provided by this invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, structures, or units, and may be electrical, mechanical, or other forms.

[0058] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0059] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0060] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for allocating GPU computing instances in a cloud server cluster, characterized in that, include: The system obtains the configuration information of the GPUs in the cluster's GPU resource pool and the instance modes of the computing instances supported by the GPUs. The computing instance modes are categorized into a first type of instance mode and a second type of instance mode. The first type of instance mode requires a complete GPU resource for execution, while the second type of instance mode, after being allocated to a GPU, requires the GPU to support at least one minimum computing instance. The GPUs in the cluster's GPU resource pool are further categorized into a first type of instance mode pre-allocated GPU group and a second type of instance mode pre-allocated GPU group. The first type of instance mode pre-allocated GPU group is specifically used to execute computing instances of the first type of instance mode, and the second type of instance mode pre-allocated GPU group is specifically used to execute computing instances of the second type of instance mode. Based on the GPU's configuration information and the instance modes of the computing instances supported by the GPU, a set of feasible starting memory blocks for any of the instance modes is pre-constructed for the GPU. When assigning a target computing instance to a selected GPU, obtain the instance mode of the target computing instance and the state of the selected GPU; Perform the following procedure to set the memory allocation for the target compute instance on the selected GPU: The algorithm iterates through all possible starting memory blocks of the instance modes of the target computing instance. For the selected GPU, it checks whether contiguous memory blocks starting from the starting memory block are available for the target computing instance, and obtains all available contiguous memory blocks as candidate memory allocation schemes. For each candidate memory allocation scheme, it obtains the GPU memory block distribution assuming the application of the candidate memory allocation scheme. Based on the memory block distribution after applying the candidate memory allocation scheme, it iterates through all feasible starting memory blocks of all instance modes supported by the GPU, and determines whether contiguous memory blocks starting from each feasible starting memory block that are larger than the memory capacity required by each instance mode are available. Based on the availability of contiguous memory blocks, it counts the number of instance modes that the GPU can satisfy after applying each candidate memory allocation scheme. The algorithm selects the candidate memory allocation scheme that maximizes the number of satisfied instance modes or the distribution-weighted scheme that maximizes the number of satisfied instance modes as the final memory allocation scheme, and then configures the target computing instance.

2. The method for allocating GPU computing instances in a cloud server cluster according to claim 1, characterized in that, Pre-construct a feasible set of starting memory blocks for the GPU, specifically including: For any instance mode, determine all possible starting memory block indices for that instance mode based on the number of computing cores and memory blocks in the GPU configuration information; For each possible starting memory block index, determine whether a contiguous memory block starting from that index can meet the memory capacity requirements of this instance mode; Add the index of the starting memory block that meets the requirements to the set of feasible starting memory blocks for this instance pattern.

3. The method for allocating GPU computing instances in a cloud server cluster according to claim 1, characterized in that, Based on the availability of contiguous memory blocks, the number of instance modes that the GPU can satisfy after applying each candidate memory allocation scheme specifically includes: The current memory block distribution status of the GPU is obtained from the GPU's state, wherein the memory block distribution status records the occupancy of each memory block; For any selected candidate memory allocation scheme, assume the distribution of GPU memory blocks after applying the candidate memory allocation scheme; Traverse all feasible starting memory blocks for all instance modes supported by the GPU. For each possible starting memory block, starting from that starting memory block, detect the number of consecutive free memory blocks. Determine whether the space capacity of the consecutive free memory blocks is greater than the memory capacity required by the target computing instance. If so, determine that consecutive memory blocks starting from that starting memory block that are greater than the memory capacity required by each instance mode are available. Then, the GPU supports the corresponding instance mode after applying the candidate memory allocation scheme. Count the total number of instance modes supported by the GPU after applying the candidate memory allocation scheme. Traverse the candidate memory allocation schemes to obtain the total number of instance modes supported by the GPU after applying each candidate memory allocation scheme.

4. The method for allocating GPU computing instances in a cloud server cluster according to claim 1, characterized in that, Based on the proportion of computing instances in the first type of instance mode, a corresponding number of GPUs are selected from the cluster's GPU resource pool to form a pre-allocated GPU group for the first type of instance mode. It is ensured that the selected GPUs come from at least a set number of servers to achieve distributed deployment. The computing instances of the first type of instance mode are then allocated to the GPUs in the pre-allocated GPU group for the first type of instance mode. Select a set number of GPUs from the remaining GPUs in the cluster's GPU resource pool to form a pre-allocated GPU group for the second type of instance mode. Ensure that the selected GPUs come from at least a set number of servers to achieve distributed deployment. Assign the computing instances of the second type of instance mode to the GPUs in the pre-allocated GPU group for the second type of instance mode.

5. The method for allocating GPU computing instances in a cloud server cluster according to claim 4, characterized in that, The first type of instance mode pre-allocated GPU group supports dynamic adjustment, including: when the pre-allocated GPU group of the first type of computing instance mode cannot meet the current computing instance's running needs due to fluctuations in the instance mode proportion, within the maximum size limit of the pre-allocated GPU group of the first type of instance mode, a new empty GPU is selected from the cluster GPU resource pool and added to the pre-allocated GPU group of the first type of instance mode, and the target computing instance of the first type of instance mode is allocated to the newly added GPU; the number of idle GPUs in the pre-allocated GPU group of the first type of instance mode is detected, and when the number of idle GPUs in the pre-allocated GPU group of the first type of instance mode exceeds a set threshold and reaches a set time, the idle GPUs are recycled to the GPU resource pool, or when the number of idle GPUs in the pre-allocated GPU group of the first type of instance mode does not exceed the set threshold but reaches the recycling period, the idle GPUs are recycled to the GPU resource pool.

6. The method for allocating GPU computing instances in a cloud server cluster according to claim 1, characterized in that, The second type of instance mode pre-allocated GPU groups support dynamic adjustment based on task requirements, including: Based on the instance mode of the compute instance, detect whether there is a GPU that supports the operation in the pre-allocated GPU group of the second type of instance mode; If it exists, the size of the pre-allocated GPU group for the second type of instance mode remains unchanged, and the target computing instance is allocated to the first detectable runnable GPU. Alternatively, all candidate GPUs that support the running of computing instances are obtained. For each candidate GPU, the number of distributed weighted instance modes satisfied after applying the target computing instance is calculated, and the candidate GPU that maximizes the number of distributed weighted instance modes satisfied by the second type of instance mode is selected as the final allocation target. If it does not exist, a new empty GPU is selected from the cluster GPU resource pool and added to the pre-allocated GPU group of the second type of instance mode, and the target computing instance is assigned to the new empty GPU. If the cluster GPU resource pool is empty, then wait. Based on resource differences or a set time period, defragmentation is performed to reclaim GPUs from the pre-allocated GPU groups in the second instance mode to the GPU resource pool.

7. The method for allocating GPU computing instances in a cloud server cluster according to claim 6, characterized in that, Based on resource differences or a set time period, defragmentation is performed to reclaim GPUs from the pre-allocated GPU group of the second-type instance mode to the GPU resource pool. This includes: detecting the GPU resource size of the pre-allocated GPU group of the second-type instance mode and the GPU resource difference required by all the second-type instance mode computing instances it executes; if the resource difference exceeds a set threshold and reaches a set duration, or if it does not exceed the set threshold but reaches the defragmentation period, then defragmentation is performed on the pre-allocated GPU group of the second-type instance mode. During the defragmentation process, the set of GPUs belonging to the pre-allocated GPU group of the second-type instance mode on any server in the cluster is determined, and it is checked whether the GPUs in the GPU set are allowed to migrate computing instances to obtain idle GPUs. If so, local GPU migration of computing instances is performed, tasks are redirected, and idle GPUs are removed from the pre-allocated GPU group of the second-type instance mode.

8. The method for allocating GPU computing instances in a cloud server cluster according to claim 7, characterized in that, During the defragmentation process of GPUs in the pre-allocated GPU group of the second type of instance mode, on each server local, the current free memory blocks and occupied memory blocks of each GPU are counted, and the weighted number of instance modes supported by the free memory blocks is counted. The ratio of free memory blocks to occupied memory blocks is multiplied by the weighted number as the evaluation value for whether the GPU should perform compute instance migration. GPUs that perform compute instance migration are selected according to the size of the evaluation value.

9. A GPU computing instance allocation device in a cloud server cluster, comprising: At least one processing unit is connected to a storage unit via a bus unit, characterized in that the storage unit stores a computer program, and the processing unit implements the GPU computing instance allocation method in any one of claims 1-8 by running the computer program stored in the storage unit.

Citation Information

Patent Citations

  • GPU cluster load balancing system based on micro-service architecture

    CN121704956A

  • Resource sharing method and device, equipment, storage medium and computer program product

    CN121833257A